Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search
TikTok USDS
- 🇺🇸 United States
- On-site
- Staff / Principal
- 3 months ago
- Machine Learning
- Node.js
- ACLS
- Python
- C++
- Java
- Kubernetes
- Slurm
- Core ML
- PyTorch
- TensorFlow
- CUDA
- InfiniBand
- Triton
- TensorRT
- MLOps
3 months ago
We are a group of applied machine learning engineers that focus on TikTok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.
What You'll Do
- Technical Leadership: Drive the technical roadmap for our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the TikTok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many businesses domains.
- Large-Scale Parallelism Architecture: Architect and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Lead the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines, establish robust SLI/SLO frameworks while maximizing hardware utilization and cluster efficiency to expedite innovation.
- Production Inference & Serving: Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
- Algorithm-Infra Co-Design: Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
- Resiliency & Fault Tolerance: Design robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
- Cross-Functional Collaboration: Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
- Security & Compliance Hardening: Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Minimum Qualifications
- Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
- 5+ years of professional software engineering experience with deep expertise in Python, C++/Java and a proven track record of designing large-scale distributed systems.
- 3+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale (managing large GPU clusters, Kubernetes, or native Slurm environments).
- Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensorflow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (InfiniBand/RoCE).
- Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
- Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
Preferred Qualifications
- Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, DeepSpeed).
- Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
- Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
- Strong flavor in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server being a huge plus.
- Active background or interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
- Excellent technical leadership skills, with the ability to mentor junior engineers, draft clear architecture designs, and communicate complex infrastructure needs to non-technical stakeholders.
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search · TikTok USDS