ES
Senior Site Reliability Engineer
EPAM Systems
๐ฒ๐ฝ Mexico | ๐ฆ๐ท Argentina | ๐จ๐ฑ Chile | ๐จ๐ด Colombia
Remote
Senior
2 days ago
- Java
- Devops
- Incident Response
- AWS
- DynamoDB
- Git
- Gradle
- Kubernetes
- Terraform
- Grafana
- New Relic
- Apache Kafka
2 days ago
We are looking for a hands-onSenior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.
Responsibilities
- Provide on-call support for Java backend identity services during business hours
- Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution
- Prepare and deploy patches to address issues in cloud infrastructure
- Improve service reliability by implementing practical changes that reduce errors and instability
- Build and refine metrics and dashboards to surface platform health and service behavior
- Monitor SLOs and propose code changes that improve SLO attainment as issues arise
- Create and improve runbooks to standardize operational response and reduce time to recovery
- Communicate incidents and operational risks clearly in writing during live response
- Collaborate closely with engineers to align operational practices with service ownership
Requirements
- 3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems
- Strong on-call support experience for production services and incident response during business hours
- Proven experience with Amazon Web Services in production environments
- Hands-on experience with Amazon DynamoDB and Amazon ElastiCache
- Strong Git skills for collaborating on operational and reliability code changes
- Solid Gradle knowledge for building and maintaining Java-based services
- Strong troubleshooting skills using logs and telemetry to identify root causes
- Clear written communication skills for documenting and reporting operational issues during incidents
- Proactive learning mindset to absorb complex information quickly and apply it under pressure
- Upper-Intermediate English proficiency (B2)
Nice to have
- Kubernetes
- Terraform
- Grafana
- New Relic
- Apache Kafka
Senior Site Reliability Engineer ยท EPAM Systems