Site Reliability Engineer, Platform Responsibility - USDS
TikTok USDS
- ๐บ๐ธ United States
- On-site
- 5 months ago
- Machine Learning
- AI
- Incident Response
- LLM APIs
- Unix
- Linux
- Prometheus
- Grafana
- Datadog
- MySQL
- Redis
- NGINX
- Kubernetes
- Docker
- Hadoop
- Flink
- Hive
- ClickHouse
5 months ago
Team Intro
The Platform Responsibility engineering team is fast growing and responsible for building machine learning models and systems to identify and defend internet abuse and fraud on our platform. Our mission is to protect billions of users and publishers across the globe every day. We embrace the state-of-the-art machine learning technologies and scale them to detect and improve trust and safety system using the tremendous amount of data generated on the platform. With the continuous efforts from our team, TikTok USDS is able to provide the best user experience and bring joy to everyone in the world.
Responsibilities
- Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
- Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
- Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
- Participate in regular on-call rotations as part of a team that provides 24 hour coverage across multiple shifts
- Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
- Practice sustainable user support, incident response, and post mortem
Minimum Qualifications
- Bachelor or above degree in computer science or a related technical discipline
- At least 1 year of industrial experience
- Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
- Demonstrated independent thinking capabilities and troubleshooting skills
- Familiar with Unix/Linux system internals, networking, and distributed systems
- Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches
Preferred Qualifications
- In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
- Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc.
Site Reliability Engineer, Platform Responsibility - USDS ยท TikTok USDS