Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Site Reliability Engineer

EPAM Systems
๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ฑ Chile | ๐Ÿ‡จ๐Ÿ‡ด Colombia
Remote
Senior
1 day ago
  • Java
  • Devops
  • Incident Response
  • AWS
  • DynamoDB
  • Git
  • Gradle
  • Kubernetes
  • Terraform
  • Grafana
  • New Relic
  • Apache Kafka
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking a hands-onSenior Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while strengthening reliability, observability, and operational readiness. You will partner closely with engineering stakeholders, contribute during on-call efforts, and drive measurable SLO improvements.

Responsibilities

  • Provide on-call support for Java backend services during business hours
  • Troubleshoot complex distributed system issues using logs and telemetry
  • Identify root causes and drive incident resolution through actionable changes
  • Prepare and deploy patches to address cloud infrastructure issues
  • Define and improve service metrics and dashboards to assess platform health
  • Improve reliability and observability posture for key services
  • Create and refine runbooks to standardize operational response
  • Track and improve SLOs through repeatable processes
  • Submit code changes that improve SLOs when errors occur
  • Communicate operational issues clearly and concisely in writing during incidents

Requirements

  • 3+ years of SRE/DevOps experience supporting production services
  • Strong on-call support experience for backend service ecosystems
  • Proven incident response skills using logs and telemetry to find root causes
  • Hands-on Amazon Web Services experience
  • Solid Amazon DynamoDB experience
  • Solid Amazon ElastiCache experience
  • Strong Git skills for contributing and reviewing code changes
  • Working Gradle knowledge in Java service environments
  • Strong observability and troubleshooting skills in distributed systems
  • Clear written communication skills for live incident updates
  • Fast learning ability to absorb information quickly and apply it under pressure
  • English proficiency: B2 Upper-Intermediate

Nice to have

  • Kubernetes experience
  • Terraform experience
  • Grafana dashboarding experience
  • New Relic monitoring experience
  • Apache Kafka experience

Senior Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote