Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

AI Solution Architect

Uvation
🇮🇳 India
Remote
Staff / Principal
11 hours ago
  • AI
  • Kubernetes
  • AI/ML
  • Ceph
  • Disaster Recovery
  • Network Security
  • CUDA
  • PyTorch
  • TensorFlow
  • JAX
  • Machine Learning
  • RAG
  • PCIe
  • InfiniBand
  • BGP
  • QoS
  • Data Architecture
  • NFS
  • MLOps
  • Secrets Management
  • AWS
  • Azure AI
  • IAM
  • FinOps
  • Zero Trust
  • RBAC
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Linux
  • TOGAF
  • CCNP
  • CCIE
  • CISSP
  • CKA
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Overview

We are seeking an experiencedAI Solution Architectto design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise inNVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.

 Key Responsibilities

  • Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
  • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
  • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
  • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
  • Design AI storage and data architectures using object storage, parallel file systems  like  Ceph, WEKA, , or equivalent platforms.
  • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
  • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
  • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
  • Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
  • Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
  • Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
  • Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.

                       

•Required Technical Skills

AI / ML Architecture

• NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.

• PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.

• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.

• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.

GPU & AI Factory Infrastructure

• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.

• NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.

• DGX/HGX/OEM GPU server architecture and lifecycle management.

• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.

High-Performance Networking

• 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris

• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.

• BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.

• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.

AI Storage & Data Architecture

• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.

• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.

• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.

• GPUDirect Storage and storage/network performance optimization.

AI Platform & Orchestration

• Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.

• HPC or other equivalent workload schedulers.

• Model serving/inference platforms and MLOps platform architecture.

• API gateways, service discovery, secrets management and platform integration.

Cloud & Hybrid Architecture

• AWS and/or Azure AI infrastructure and security services.

• Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.

• Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.

Security & Governance

• Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.

• GPU, DPU, container, Kubernetes, firmware and supply-chain security.

• Encryption at rest/in transit, secrets management, audit logging and compliance controls.

• AI-specific risks including data/model protection, tenant isolation and secure model access.

Observability & Reliability

• Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.

• Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.

• High availability, backup/restore, disaster recovery, business continuity and failure-domain design.

• Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.

Architecture Deliverables

• AI Factory reference architecture and solution blueprints

• High-Level Design (HLD) and Low-Level Design (LLD)

• Network, compute, GPU and storage architecture diagrams

• Capacity, performance and scalability models

• Technology evaluation and vendor comparison documents

• Security architecture and threat-model inputs

• Bill of Materials (BOM) and infrastructure sizing

• Migration/deployment strategy and implementation roadmap

• Operational readiness checklist, runbooks and acceptance criteria

Experience & Qualifications

• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.

• Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.

• Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.

• Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.

Preferred Certifications

• NVIDIA certifications or equivalent GPU/AI infrastructure credentials

• AWS Solutions Architect / Azure Solutions Architect

• TOGAF or equivalent enterprise architecture certification

• CCNP/CCIE or equivalent networking certification

• CISSP or equivalent security certification

• Kubernetes certifications such as CKA/CKAD

• Red Hat / Linux certifications

AI Solution Architect · Uvation

Auto apply with Likeremote