
AI Solution Architect
- AI
- Kubernetes
- AI/ML
- Ceph
- Disaster Recovery
- Network Security
- CUDA
- PyTorch
- TensorFlow
- JAX
- Machine Learning
- RAG
- PCIe
- InfiniBand
- BGP
- QoS
- Data Architecture
- NFS
- MLOps
- Secrets Management
- AWS
- Azure AI
- IAM
- FinOps
- Zero Trust
- RBAC
- Prometheus
- Grafana
- OpenTelemetry
- Linux
- TOGAF
- CCNP
- CCIE
- CISSP
- CKA
Job Overview
We are seeking an experiencedAI Solution Architectto design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise inNVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.
 Key Responsibilities
- Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
- Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
- Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
- Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
- Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.
- Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
- Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
- Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
- Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
- Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
- Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
- Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.
                     Â
•Required Technical Skills
AI / ML Architecture
• NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.
• PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.
• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
GPU & AI Factory Infrastructure
• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.
• NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
• DGX/HGX/OEM GPU server architecture and lifecycle management.
• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
High-Performance Networking
• 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris
• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
• BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.
• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
AI Storage & Data Architecture
• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
• GPUDirect Storage and storage/network performance optimization.
AI Platform & Orchestration
• Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.
• HPC or other equivalent workload schedulers.
• Model serving/inference platforms and MLOps platform architecture.
• API gateways, service discovery, secrets management and platform integration.
Cloud & Hybrid Architecture
• AWS and/or Azure AI infrastructure and security services.
• Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.
• Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.
Security & Governance
• Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.
• GPU, DPU, container, Kubernetes, firmware and supply-chain security.
• Encryption at rest/in transit, secrets management, audit logging and compliance controls.
• AI-specific risks including data/model protection, tenant isolation and secure model access.
Observability & Reliability
• Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
• Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
• High availability, backup/restore, disaster recovery, business continuity and failure-domain design.
• Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.
Architecture Deliverables
• AI Factory reference architecture and solution blueprints
• High-Level Design (HLD) and Low-Level Design (LLD)
• Network, compute, GPU and storage architecture diagrams
• Capacity, performance and scalability models
• Technology evaluation and vendor comparison documents
• Security architecture and threat-model inputs
• Bill of Materials (BOM) and infrastructure sizing
• Migration/deployment strategy and implementation roadmap
• Operational readiness checklist, runbooks and acceptance criteria
Experience & Qualifications
• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
• Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
• Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
• Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
Preferred Certifications
• NVIDIA certifications or equivalent GPU/AI infrastructure credentials
• AWS Solutions Architect / Azure Solutions Architect
• TOGAF or equivalent enterprise architecture certification
• CCNP/CCIE or equivalent networking certification
• CISSP or equivalent security certification
• Kubernetes certifications such as CKA/CKAD
• Red Hat / Linux certifications
AI Solution Architect · Uvation