Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead Site Reliability Engineer

EPAM Systems
๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ฑ Chile | ๐Ÿ‡จ๐Ÿ‡ด Colombia
Remote
Staff / Principal
2 days ago
  • Java
  • Devops
  • AWS
  • DynamoDB
  • Git
  • Gradle
  • Incident Response
  • Kubernetes
  • Terraform
  • Grafana
  • Apache Kafka
  • New Relic
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for a hands-onLead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex production issues using logs and telemetry to identify root causes
  • Prepare and deploy patches to address issues in cloud infrastructure
  • Implement reliability improvements for key identity services through practical code and configuration changes
  • Build and refine metrics and dashboards to enable rapid assessment of platform health
  • Monitor SLOs across backend services and drive remediation when error rates increase
  • Create and improve runbooks to standardize operational responses across services

Requirements

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
  • Strong experience with Amazon Web Services in production environments
  • Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
  • Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
  • Hands-on experience with Git-based workflows
  • Hands-on experience with Gradle in Java service environments
  • Leadership skills to guide reliability improvements and support operational decision-making
  • Incident response skills to communicate operational issues clearly and concisely in writing
  • Fast learning ability to absorb information quickly and apply it during on-call support
  • SLO management skills to track, evaluate, and improve reliability through repeatable processes
  • English proficiency: B2 (Upper-Intermediate)

Nice to have

  • Kubernetes
  • Terraform
  • Grafana
  • Apache Kafka
  • New Relic

Lead Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote