Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TikTok USDS logo

Senior Machine Learning Engineer - LLM Evaluation & Agent Systems

TikTok USDS
  • 🇺🇸 United States
  • On-site
  • Senior
  • 1 week ago
  • AI
  • Core ML
  • System Design
  • Machine Learning
  • RAG
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About the Team

The TikTok U.S. Data Security (USDS) team is responsible for the security, integrity, and compliance of TikTok’s data within the United States. The Data Platform Team builds a scalable, reliable, cost-efficient infrastructure that ensures data integrity and empowers every business unit to win through self-service tools and scientific insights.

About the Role

We are seeking a Senior Machine Learning Engineer to join our team as a technical lead, driving the architecture and execution of our next-generation AI Risk & Compliance Intelligence Platform and Intelligent Data Agent Ecosystem.

In this role, you will be instrumental in bridging the gap between standard prompt engineering and native, model-driven reasoning frameworks. You will tackle complex ML challenges around offline/online evaluation, synthetic data generation for small-sample scenarios, and automated agent skill optimizations across Data & Engineering workflows (DA/DS/DE/SRE).

Key Responsibilities

  • Design, build, and scale robust LLM evaluation architecture, including automated offline datasets, metric suites, and benchmarking pipelines to measure prompt and model performance.
  • Innovate on LLM-driven synthetic data generation methodologies to tackle data scarcity bottlenecks and systematically evaluate long-tail/edge-case risk scenarios.
  • Advance beyond criteria-based prompt engineering by extracting deep risk patterns directly from complex case data to build native, ML-driven decision models and reasoning abstractions.
  • Architect, evaluate, and continuously optimize modular agent skills and tools tailored for Data Science, Analytics, Data Engineering, and SRE workflows to maximize tool-use accuracy, execution efficiency, and task success rates.
  • Implement advanced evaluation methodologies (such as LLM-as-a-judge and multi-agent benchmarking) across both risk intelligence and agentic ecosystems.
  • Provide technical leadership by establishing core ML system design patterns, benchmarking standards, and engineering best practices, while translating complex operational requirements into scalable machine learning solutions across cross-functional partnerships.

Minimum Qualifications

  • Experience: 5+ years of software/ML engineering experience with a strong track record of designing, building, and deploying ML/LLM systems in production.
  • Generative AI & LLMs: Deep hands-on expertise with LLM architectures, fine-tuning, RAG, prompt tuning, and agentic frameworks.
  • Evaluation & Data Bootstrapping: Demonstrated experience building robust LLM evaluation frameworks (LLM-as-a-judge, automated benchmark pipelines) and tackling data scarcity/imbalance problems.
  • Core Engineering: Strong coding skills alongside production experience with scalable data processing pipelines.

Preferred Qualifications

  • Prior domain experience in Risk & Compliance Intelligence and Automated Decision Systems.
  • Hands-on experience with synthetic dataset generation or bootstrapping models for cold-start scenarios.
  • Experience building AI tools or agents for technical personas (Data Analysts, Data Scientists, Data Engineers, Data SRE).
  • Demonstrated ability to serve as a founding technical lead on ambiguous, high-impact initiative pods.

Senior Machine Learning Engineer - LLM Evaluation & Agent Systems · TikTok USDS

Auto apply with Likeremote