L
Associate Principal - Architecture
LTM
- Location not stated
- Staff / Principal
- 1 day ago
- AI
- Kubernetes
- MLOps
- Bootstrap
- IaC
- GitOps
- CI/CD
- FinOps
- IAM
- KMS
- VPC
- AWS
- Azure
- GCP
- Linux
- EKS
- AKS
- GKE
- OpenShift
- RHEL
- CentOS
- Ubuntu
- Terraform
- Disaster Recovery
- Nim
- Vertex AI
- Azure ML
1 day ago
AI Platform Architect – GPU, Kubernetes & AI Infrastructure
Experience: 12- 16 Years
Role: AI Platform Architect
Domain: AI Infrastructure, Kubernetes, GPU Platforms, MLOps/GenAIOps
Employment Type: Full-Time
Key Responsibilities
- Design enterprise AI platforms, GPU-enabled Kubernetes environments, and cloud/on-prem landing zones.
- Define platform blueprints, site bootstrap automation, Infrastructure-as-Code, and GitOps operating models.
- Architect multi-tenant AI platforms, tenant onboarding frameworks, and lifecycle management processes.
- Design AI model serving, inference, training, autoscaling, GPU scheduling, and resource-sharing architectures.
- Establish MLOps/GenAIOps frameworks including model registry, CI/CD, observability, and FinOps integration.
- Define HA/DR strategies, cross-region resiliency, failover mechanisms, and recovery workflows.
- Lead architecture reviews, design workshops, and technical governance across engineering teams.
- Act as design authority for AI, Kubernetes, and GPU platform architecture decisions.
Must-Have Skills
Platform & Cloud Architecture
- 12+ years of IT experience with 5+ years in cloud/platform architecture.
- Strong experience designing large-scale cloud and hybrid infrastructure platforms.
- Expertise in landing zones, IAM/KMS, networking, VPC/VNet, hybrid connectivity, and multi-account architectures.
- Professional cloud certification (AWS/Azure/GCP).
Kubernetes & Linux
- Deep Kubernetes architecture and operational experience (EKS, AKS, GKE, OpenShift, On-Prem K8s).
- Strong expertise in multi-tenancy, workload isolation, cluster design, and platform operations.
- Advanced Linux administration, hardening, networking, and troubleshooting (RHEL/CentOS/Ubuntu).
Infrastructure Automation
- Hands-on experience with Terraform or equivalent Infrastructure-as-Code tools.
- Strong knowledge of GitOps, automation pipelines, and platform engineering practices.
AI Platform & GPU Infrastructure
- Experience designing AI training and inference platforms.
- Knowledge of GPU scheduling, workload orchestration, model serving, and inference scaling.
- Exposure to multi-tenant AI platform security, quota management, and workload isolation.
High Availability & Resiliency
- Expertise in HA/DR, multi-region architecture, disaster recovery, failover, and RTO/RPO planning.
Stakeholder Management
- Strong customer-facing experience conducting architecture workshops and executive design reviews.
- Ability to translate business requirements into scalable technical solutions.
Preferred Skills
- NVIDIA AI Enterprise (NVAIE), NIM, Dynamo, RunAI.
- GPU sharing technologies (MIG, Fractional GPU).
- Rafay, Spectro Cloud, vCluster, OPA, Kyverno.
- SageMaker, Vertex AI, Azure ML.
- Telco, Edge, 5G MEC, AWS Outposts, Azure Local.
Associate Principal - Architecture · LTM