Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
K

Senior Site Reliability Engineer

Kody
๐Ÿ‡จ๐Ÿ‡ณ China | ๐Ÿ‡ญ๐Ÿ‡ฐ Hong Kong
On-site
Senior
1 week ago
  • Incident Response
  • Incident Management
  • triage
  • Kubernetes
  • PCI DSS
  • Devops
  • AWS
  • EKS
  • Terraform
  • PostgreSQL
  • Redis
  • Linux
  • Datadog
  • Prometheus
  • Grafana
  • Disaster Recovery
  • Microservices
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Summary

Kody is seeking aSenior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based inShenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities

  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

Qualifications & Requirements

  • Experience:8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise inAWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based inHong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Leadership & Operational Excellence

  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

- Competitive package

- A dynamic and innovative team

- Collaborative, inclusive working environment where your contributions are recognized

Senior Site Reliability Engineer ยท Kody

Auto apply with Likeremote