ES
Lead Site Reliability Engineer
EPAM Systems
- ๐ง๐ท Brazil | ๐ฒ๐ฝ Mexico | ๐ฆ๐ท Argentina | ๐จ๐ด Colombia
- Remote
- Staff / Principal
- 14 hours ago
- Devops
- CI/CD
- GitLab
- Python
- Kubernetes
- IAM
- Incident Response
- AWS
- Azure
- AI
14 hours ago
We are looking for aLead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate safe delivery of change in a fast-moving environment. You will elevate DevOps maturity across tooling, processes, and engineering practices while solving complex reliability challenges.
Responsibilities
- Own reliability outcomes for critical infrastructure and set SRE engineering standards
- Design and implement scalable DevOps tooling and processes that improve delivery speed and safety
- Build and maintain CI/CD workflows and source control practices using GitLab where appropriate
- Automate infrastructure operations with Python to reduce toil and improve consistency
- Improve Kubernetes-based workflows to support dependable deployments and runtime stability
- Diagnose and resolve business-critical incidents during on-call shifts with urgency and rigor
- Evaluate and harden infrastructure domains such as networking, compute, security, IAM, and configuration automation
- Drive enterprise-grade release management practices across environments
- Partner with stakeholders to translate reliability needs into actionable engineering plans
- Promote long-term engineering solutions over quick fixes through root-cause analysis and follow-up actions
Requirements
- 5+ years of site reliability engineering experience in production environments
- 5+ years of cloud platform experience with a leading provider
- Proven leadership skills to drive reliability improvements and mentor engineers
- Enterprise-scale release management experience across complex systems
- Strong CI/CD knowledge covering pipelines, source control, and automation practices
- Advanced Python programming skills for automation and engineering solutions
- Hands-on Kubernetes experience as a developer in delivery workflows
- Strong infrastructure fundamentals across networking, compute, security, and IAM
- Excellent analytical skills for strategic thinking and complex problem solving
- Effective incident response skills, including on-call ownership for critical issues
- English proficiency: B2 Upper-Intermediate
Nice to have
- Amazon Web Services expertise in production environments
- Microsoft Azure expertise in production environments
- AI Architecture experience for reliability-aware AI platform design
- AI Solution Engineering experience supporting clients with scalable implementations
- Gen AI Solutions Development experience in enterprise settings
Lead Site Reliability Engineer ยท EPAM Systems