Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TikTok USDS logo

Senior Machine Learning Infrastructure Engineer, Recommendations and Search

TikTok USDS
  • 🇺🇸 United States
  • On-site
  • Senior
  • 1 month ago
  • Machine Learning
  • Node.js
  • ACLS
  • Python
  • C++
  • Java
  • Core ML
  • PyTorch
  • TensorFlow
  • CUDA
  • InfiniBand
  • Triton
  • TensorRT
  • MLOps
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About the Team

We are a group of applied machine learning engineers that focus on TikTok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.

What You'll Do

  • Technical Execution & System Ownership: Contribute to the technical roadmap and hands-on implementation of our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the TikTok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many business domains.
  • Large-Scale Parallelism Architecture: Help design and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Participate in the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines to achieve high hardware utilization and cluster efficiency under established SLI/SLO frameworks.
  • Production Inference & Serving: Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
  • Algorithm-Infra Co-Design: Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
  • Resiliency & Fault Tolerance: Build robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
  • Cross-Functional Collaboration: Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
  • Security & Compliance Hardening: Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.

Minimum Qualifications

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 4+ years of professional software engineering experience with deep expertise in Python, C++/Java
  • 2+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale
  • Deep technical familiarity with the internals of core ML frameworks (PyTorch / TensorFlow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (InfiniBand/RoCE).
  • Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.

Preferred Qualifications

  • Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, DeepSpeed).
  • Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
  • Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
  • Strong flavor in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server being a huge plus.
  • Active background or interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
  • Strong communication skills, with the ability to collaborate effectively across team boundaries, draft clear technical design docs, and mentor team members.

Senior Machine Learning Infrastructure Engineer, Recommendations and Search · TikTok USDS

Auto apply with Likeremote