ES
Lead Site Reliability Engineer
EPAM Systems
๐ฒ๐ฝ Mexico | ๐ฆ๐ท Argentina | ๐จ๐ฑ Chile | ๐จ๐ด Colombia
Remote
Staff / Principal
2 days ago
- Java
- Devops
- AWS
- DynamoDB
- Git
- Gradle
- Incident Response
- Kubernetes
- Terraform
- Grafana
- Apache Kafka
- New Relic
2 days ago
We are looking for a hands-onLead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.
Responsibilities
- Provide on-call support for Java backend identity services during business hours
- Troubleshoot complex production issues using logs and telemetry to identify root causes
- Prepare and deploy patches to address issues in cloud infrastructure
- Implement reliability improvements for key identity services through practical code and configuration changes
- Build and refine metrics and dashboards to enable rapid assessment of platform health
- Monitor SLOs across backend services and drive remediation when error rates increase
- Create and improve runbooks to standardize operational responses across services
Requirements
- 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
- Strong experience with Amazon Web Services in production environments
- Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
- Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
- Hands-on experience with Git-based workflows
- Hands-on experience with Gradle in Java service environments
- Leadership skills to guide reliability improvements and support operational decision-making
- Incident response skills to communicate operational issues clearly and concisely in writing
- Fast learning ability to absorb information quickly and apply it during on-call support
- SLO management skills to track, evaluate, and improve reliability through repeatable processes
- English proficiency: B2 (Upper-Intermediate)
Nice to have
- Kubernetes
- Terraform
- Grafana
- Apache Kafka
- New Relic
Lead Site Reliability Engineer ยท EPAM Systems