Engineering Manager, ML Performance Optimization
- 🇺🇸 United States
- Hybrid
- Manager or above
- 4 hours ago
- $276,000 – $343,000 / year
- CUDA
- Triton
- PyTorch
- JAX
- TensorRT
- Ray
- Machine Learning
- Health insurance
- Equity
In this role, you will:
Vision: Develop and execute a strategic vision and roadmap for ML Training and Inference Performance Optimization, ensuring scalability, reliability, and performance to support autonomous driving.
Technical acumen: Lead the design, implementation, and operation of a robust and efficient ML platform to enable the training, validation, serving, optimization and monitoring of ML models.
ML Performance Optimization: Drive end-to-end performance optimization for large-scale model training and inference, including distributed training efficiency, GPU utilization, memory and communication optimization, model compression (quantization, pruning, distillation), and low-latency on-vehicle inference that meets strict real-time and compute budgets.
Hiring: Attract, hire, and inspire a diverse world-class engineering team, fostering a culture of innovation, collaboration, and excellence.
Partnership: Collaborate closely with cross-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions.
Mentorship: Enable engineers on the team to grow their careers by providing the right opportunities and clear, timely feedback.
Qualifications
- 8+ years of relevant experience, including 3+ years of management experience managing engineers.
- Strong technical background in ML performance optimization, such as distributed training strategies (data, tensor, pipeline parallelism, FSDP/ZeRO), mixed-precision training, kernel-level optimization (CUDA, Triton), compiler stacks (torch.compile, XLA, TVM), quantization, and profiling/benchmarking across GPU and embedded accelerators.
- Experience building user-friendly ML Infrastructure that enabled large-scale model training and high-throughput, low-latency serving use cases.
- Experience with training frameworks like PyTorch, JAX, etc., leveraging GPUs for distributed model training.
- Experience with GPU-accelerated inference using TensorRT, Ray Serve, or similar frameworks.
- Proven track record of extensive cross-functional collaboration, partnering with research, product, hardware, and platform teams to align priorities, influence technical direction, and deliver measurable performance improvements across organizational boundaries.
Follow us on LinkedIn
AccommodationsIf you need an accommodation to participate in the application or interview process please reach out to accommodations@zoox.com or your assigned recruiter.
A Final Note:You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.
Engineering Manager, ML Performance Optimization · Zoox