Data Site Reliability Engineer, Video Platform
- ๐ฆ๐บ Australia
- On-site
- Mid level
- 3 months ago
- SQL
- Java
- Python
- Scala
- Hadoop
- Flink
- Yarn
- NoSQL
- Airflow
- AWS
- Azure
Team Intro
TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding services. As part of the USDS, the Video Platform team is responsible for building the next generation video processing platform which provides excellent experiences for billions of users around the world.
The USDS Video Platform team is seeking an experienced Site Reliability Engineer to help us continue improving TikTok's video system. If you are passionate about ensuring software reliability, love problem-solving, and are prepared for exciting challenges, we would like you on our team.
Responsibilities
- Manage day-to-day operations of data service, realtime/batch data pipelines, such as Service Level Agreement management, pipeline deployment, performance tuning and troubleshooting
- Proactively monitor and troubleshoot data pipelines and systems for performance issues, errors, or anomalies
- Create tools, build alarms and dashboards, drive internal process improvements, and automation to monitor and improve data engineering operations
- Improve systems reliability, efficiency, and velocity through scaling, optimization of both resources and data processing workflows, potentially refactoring code or implementing new solutions
- Develop and deploy new reliable and scalable data pipelines and infrastructure components as required by business needs
- Work closely with data engineering and various vertical teams within the Video Architecture platform
Minimum Qualifications
- Bachelor's in Computer Science or a related technical background involving software/system engineering, or equivalent working experience
- Good programming experience with SQL and at least one of the following languages: Java, Python, Go, or Scala
- Experience in data engineering, with a focus on data systems reliability, scalability, performance and capacity management
- Solid experience with big data technologies (e.g., Hadoop, Spark, Flink, YARN) and databases (SQL, NoSQL)
- Knowledge of data pipeline and workflow management tools (e.g., Airflow, Luigi)
- Experience in building data solutions with AWS, Google, Azure and other cloud services is a plus
Preferred Qualifications
- Demonstrated independent thinking capabilities and troubleshooting skills in large scale distributed systems
- Good communication and coordination skills
Data Site Reliability Engineer, Video Platform ยท TikTok USDS