Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Site Reliability Engineer (SRE)

EPAM Systems
πŸ‡¦πŸ‡² Armenia | πŸ‡°πŸ‡¬ Kyrgyzstan | πŸ‡°πŸ‡Ώ Kazakhstan | πŸ‡ΊπŸ‡Έ United States | πŸ‡ΊπŸ‡Ώ Uzbekistan
Remote
Senior
1 week ago
  • Grafana
  • AWS
  • Kubernetes
  • EKS
  • Incident Response
  • Load Testing
  • IaC
  • AIOps
  • Devops
  • Prometheus
  • Loki
  • OpenTelemetry
  • Python
  • Bash
  • Terraform
  • CI/CD
  • Incident Management
  • triage
  • k6
  • JMeter
  • Datadog
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for aSenior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.

Responsibilities

  • Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks
  • Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality
  • Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms
  • Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling
  • Automate operational toil through scripting and infrastructure-as-code
  • Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort
  • Lead incident response practices including on-call readiness and blameless post-mortems
  • Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements

Requirements

  • 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
  • Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
  • Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
  • Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
  • Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
  • Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
  • A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
  • Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
  • English Level: B2+ (Upper-Intermediate) or higher

Nice to have

  • Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
  • Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
  • Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
  • Experience with Datadog or similar enterprise observability platforms
  • Background in evangelizing best practices and setting standards across engineering teams
  • Exposure to programmatic advertising or adtech platforms

Senior Site Reliability Engineer (SRE) Β· EPAM Systems

Auto apply with Likeremote