Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TikTok USDS logo

Major Incident Manager, Incident Management -TikTok USDS

TikTok USDS
  • 🇺🇸 United States
  • On-site
  • Manager or above
  • 2 months ago
  • Incident Management
  • Incident Response
  • AWS
  • GCP
  • Azure
  • Microservices
  • Kubernetes
  • CI/CD
  • Grafana
  • Splunk
  • Prometheus
  • New Relic
  • Bash
  • Python
  • ITIL
  • Devops
  • Ansible
  • PagerDuty
  • Opsgenie
  • Jira Service Management
  • SQL
  • Tableau
  • Data Visualization
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

The USDS JV Incident Management Team (IMT) is a critical pillar within the US Tech and Product organization, dedicated to ensuring the resilience and reliability of TikTok’s U.S. Data Security infrastructure. As a security-first division, we focus on providing specialized oversight and protection for U.S. user data and the platforms that support them.

While global teams monitor overall service health, the IMT is uniquely positioned to manage the "blast radius" within the USDS environment. We act as the bridge between technical engineering teams (SRE, Infrastructure, Platform) and business stakeholders, ensuring that every major incident is handled with the precision and urgency required by our unique operating environment.

As a Major Incident Manager, you will be at the forefront of protecting the TikTok experience for millions of users. You will join a cohesive, proficient team committed to upholding the highest standards of professionalism and technical expertise, directly contributing to the safety and security of the U.S. digital ecosystem.

Responsibilities

  • Incident Command: Lead the end-to-end response for high-severity incidents, driving rapid restoration of service while coordinating across cross-functional engineering teams.
  • Stakeholder Communication: Craft and deliver clear, timely, and accurate communication updates to internal stakeholders, executive leadership, and customer-facing teams during critical events.
  • Post-Incident Reviews: Facilitate blameless post-mortems, ensuring root causes are thoroughly identified, actionable remediation items are tracked, and lessons learned are shared globally.
  • Dashboarding & Visibility: Create and maintain real-time operational dashboards that provide high visibility into system health, incident trends, and key reliability metrics (MTTD/MTTR).
  • Continuous Improvement: Analyze incident data and operational metrics to identify systemic trends, driving initiatives to reduce Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR).
  • Team Process Evolution: Manage and optimize the incident management rotation, refining playbooks, alerting thresholds, and escalation pathways for the SRE org.
  • Alerting & Monitoring Improvements: Partner with SRE and development teams to continuously refine alert thresholds, reduce alert fatigue, and ensure high-severity pages are highly actionable.
  • Automation Engineering: Design and build automated workflows to streamline incident response, such as automated stakeholder communications, auto-remediation scripts, and tool integrations.
  • Oncall Support: The Incident Management team operates 24x5x365 with on-call rotation during weekends. This position is for the Sunday to Thursday shift, from 5 PM PT to 2 AM PT.

Minimum Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or a related technical field (or equivalent practical experience) with 4+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment.
  • Strong understanding of cloud computing concepts (AWS, GCP, or Azure) and modern infrastructure architectures (microservices, Kubernetes, CI/CD pipelines).
  • Hands-on experience configuring and refining alerts and monitoring templates using industry-standard tools (e.g., Grafana, Splunk, Prometheus, New Relic).
  • Ability to write scripts (e.g., Bash, Python) or use low-code integration tools to automate repetitive incident management tasks and notification pipelines.
  • Exceptional verbal and written communication skills, with a track record of translating deeply technical issues into clear business impact summaries for leadership.
  • Strong working knowledge of ITIL incident management frameworks or modern DevOps/SRE incident response practices.

Preferred Qualifications

  • 5+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment.
  • Deep alignment with Google SRE principles, including Error Budgets, SLA/SLO/SLI management, and automation-first mindsets.
  • Proven experience implementing ChatOps, auto-remediation workflows, or Event-Driven Ansible/Runbook automation to programmatically resolve common alerts.
  • Advanced proficiency with modern incident orchestration platforms like PagerDuty, Opsgenie, JIRA Service Management, and Slack integrations.
  • Experience using SQL, Tableau, or similar data visualization tools to build operational dashboards and report on organizational reliability metrics.
  • Relevant industry certifications such as ITIL v4, AWS/GCP Certified Professional, or Certified Incident Commander.
  • Experience coaching and training engineering teams on best practices for on-call hygiene and conducting constructive post-mortems.

Major Incident Manager, Incident Management -TikTok USDS · TikTok USDS

Auto apply with Likeremote