Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ByteDance logo

Site Reliability Engineer - AI Application

ByteDance
  • πŸ‡ΈπŸ‡¬ Singapore
  • On-site
  • 9 months ago
  • AI
  • Linux
  • Python
  • Java
  • Ansible
  • AWS
  • GCP
  • NGINX
  • Kubernetes
  • Docker
  • OpenStack
  • Hadoop
  • Flink
  • System Design
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About the team

We are an AI-driven search and recommendation team focused on building innovative, scalable products for global users.

Responsibilities

  • Ensure the reliability and normal operation of multiple core systems related to Viking Team's Big data and online services, while focusing on system capacity planning and stability assurance;
  • Enhance system visibility by monitoring the availability and performance metrics of system components, helping development teams quickly locate faults, and especially ensuring operation of critical links such as AI search/vector databases;
  • Improve the reliability, scalability, and Performance optimization of services to ensure the achievement of the core system SLA;
  • Participated in the design and implementation of the automation platform, ensuring the rapid iteration and efficient operation and maintenance of large-scale online Viking clusters and AI search-related clusters;
  • Combining with the usage scenarios of AI Search/Viking business, in-depth optimization of service governance practices, including but not limited to analysis of performance bottlenecks in key AI Search/Viking links, business problem location and troubleshooting, promoting the transformation and upgrading of the system's high-availability architecture, and those familiar with Viking-related technologies are preferred to participate in core optimization work.

Minimum Qualifications

  • Bachelor's degree or above, majoring in computer-related fields, with more than five years of relevant work experience;
  • Has a solid foundation in computer software knowledge, and understands the relevant principles of Linux operating systems, storage, network IO, etc.
  • Familiar with at least one programming language (such as Python/Go/Java/Shell/Ansible), with moderate development capabilities, and placing more emphasis on operations and maintenance practices and problem-solving abilities;
  • Understand at least one type of knowledge related to cloud infrastructure such as AWS/Volcano Engine/Aliyun/GCP; those with experience in computing/distributed systems are preferred (e.g., Nginx/Kubernetes/Docker/OpenStack/Hadoop/Spark/Flink, etc.);

Preferred Qualifications

  • Familiar with algorithmic thinking, good data structure and system design capabilities
  • Have certain understanding of AI Cloud, large model-related Search Suggestion, and Recommender system.

Site Reliability Engineer - AI Application Β· ByteDance

Auto apply with Likeremote