ES
Senior Site Reliability Engineer
EPAM Systems
๐ฒ๐ฝ Mexico | ๐ฆ๐ท Argentina | ๐จ๐ฑ Chile | ๐จ๐ด Colombia
Remote
Senior
1 day ago
- Java
- Devops
- Incident Response
- AWS
- DynamoDB
- Git
- Gradle
- Kubernetes
- Terraform
- Grafana
- New Relic
- Apache Kafka
1 day ago
We are seeking a hands-onSenior Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while strengthening reliability, observability, and operational readiness. You will partner closely with engineering stakeholders, contribute during on-call efforts, and drive measurable SLO improvements.
Responsibilities
- Provide on-call support for Java backend services during business hours
- Troubleshoot complex distributed system issues using logs and telemetry
- Identify root causes and drive incident resolution through actionable changes
- Prepare and deploy patches to address cloud infrastructure issues
- Define and improve service metrics and dashboards to assess platform health
- Improve reliability and observability posture for key services
- Create and refine runbooks to standardize operational response
- Track and improve SLOs through repeatable processes
- Submit code changes that improve SLOs when errors occur
- Communicate operational issues clearly and concisely in writing during incidents
Requirements
- 3+ years of SRE/DevOps experience supporting production services
- Strong on-call support experience for backend service ecosystems
- Proven incident response skills using logs and telemetry to find root causes
- Hands-on Amazon Web Services experience
- Solid Amazon DynamoDB experience
- Solid Amazon ElastiCache experience
- Strong Git skills for contributing and reviewing code changes
- Working Gradle knowledge in Java service environments
- Strong observability and troubleshooting skills in distributed systems
- Clear written communication skills for live incident updates
- Fast learning ability to absorb information quickly and apply it under pressure
- English proficiency: B2 Upper-Intermediate
Nice to have
- Kubernetes experience
- Terraform experience
- Grafana dashboarding experience
- New Relic monitoring experience
- Apache Kafka experience
Senior Site Reliability Engineer ยท EPAM Systems