L
Associate Principal - Cloud Engineering
LTM
๐บ๐ธ United States
Remote
Staff / Principal
1 month ago
- AI
- EKS
- Nim
- Microservices
- CUDA
- Triton
- vLLM
- Kubernetes
- RAG
1 month ago
100% Remote
Max Salary $164K+Benefits
Role Summary
The NVIDIA GPU Platform Engineer builds and operates the GPU and inferenceserving stack NVIDIA AI Enterprise on EKS NIM microservices and NVIDIA Dynamo serving and owns GPU sharing scheduling integration and serving performance Onshore placement supports handson GPUedge coordination and close work with the architecture team
Key Responsibilities
- Deploy and operate NVIDIA AI Enterprise NVAIE on EKS GPU Operator drivers CUDA runtime and DCGM
- Deploy and operate NIM microservices and NVIDIA Dynamo serving exposing the OpenAIcompatible API chatcompletions embeddings streaming batch
- Configure GPU sharing MIG where the SKU supports it and RunAI fractional GPU as the isolation baseline and clearly distinguish hardware vs software isolation
- Integrate RunAI for GPU scheduling quota and preemption and KEDA for replica autoscaling including scaletozero
- Instrument GPU telemetry via DCGM and drive inference performance validation and SLAtier binding
- Support the modelonboarding NIM packaging and inferencedeploy pipelines
- Tune GPUserving performance and troubleshoot CUDA driver and scheduling issues
- MustHave Skills Experience
- 8 years in infrastructureML engineering with handson NVIDIA GPU operations
- NVIDIA GPU stack drivers CUDA GPU Operator and DCGM
- Model serving Triton NVIDIA Dynamo NIM TensorRTLLM andor vLLM
- Kubernetes GPU workloads and device plugins
- GPU scheduling RunAI or equivalent and GPU partitioning MIG fractional GPU
- Autoscaling KEDA and OpenAIcompatible inference API patterns
- NicetoHave Skills
- NVAIE on AWSEKS and edge GPU Outposts G7
- LLM RAG serving and vector databases
- Inference performance benchmarking and optimization
โ
Associate Principal - Cloud Engineering ยท LTM