Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Site Reliability Engineer

EPAM Systems
๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ฑ Chile | ๐Ÿ‡จ๐Ÿ‡ด Colombia
Remote
Senior
2 days ago
  • Java
  • Devops
  • Incident Response
  • AWS
  • DynamoDB
  • Git
  • Gradle
  • Kubernetes
  • Terraform
  • Grafana
  • New Relic
  • Apache Kafka
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for a hands-onSenior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution
  • Prepare and deploy patches to address issues in cloud infrastructure
  • Improve service reliability by implementing practical changes that reduce errors and instability
  • Build and refine metrics and dashboards to surface platform health and service behavior
  • Monitor SLOs and propose code changes that improve SLO attainment as issues arise
  • Create and improve runbooks to standardize operational response and reduce time to recovery
  • Communicate incidents and operational risks clearly in writing during live response
  • Collaborate closely with engineers to align operational practices with service ownership

Requirements

  • 3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems
  • Strong on-call support experience for production services and incident response during business hours
  • Proven experience with Amazon Web Services in production environments
  • Hands-on experience with Amazon DynamoDB and Amazon ElastiCache
  • Strong Git skills for collaborating on operational and reliability code changes
  • Solid Gradle knowledge for building and maintaining Java-based services
  • Strong troubleshooting skills using logs and telemetry to identify root causes
  • Clear written communication skills for documenting and reporting operational issues during incidents
  • Proactive learning mindset to absorb complex information quickly and apply it under pressure
  • Upper-Intermediate English proficiency (B2)

Nice to have

  • Kubernetes
  • Terraform
  • Grafana
  • New Relic
  • Apache Kafka

Senior Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote