DL
Azure SRE lead engineers
Diverse Lynx India
đźđł India
On-site
Staff / Principal
2 months ago
- Incident Management
- Azure
- Kubernetes
- OpenShift
- Datadog
- Dynatrace
- Splunk
- Jenkins
- Ansible
- CI/CD
- Python
- RabbitMQ
- Java
- Node.js
- Incident Response
- AIOps
2 months ago
We are seeking a Site Reliability Engineer (SRE) with 7+ years of experience to support and enhance the reliability, availability, and performance of critical banking systems at Truist. The role requires strong handsâon expertise in cloudânative platforms, observability, automation, and incident management, with a focus on reliability engineering and operational excellence
Required Technical Skills
Cloud & Infrastructure
- Microsoft Azure
- Kubernetes
- OpenShift
Observability & Monitoring
- Datadog
- Dynatrace / AppDynamics
- Splunk
- Jenkins
- Ansible
Automation & CI/CD
- Python
- Kafka
- RabbitMQ
- Exposure to Java and Node.js
Production Support & Incident Management
- Strong experience handling major incidents
- Production support in highâavailability, missionâcritical environments
- Root cause analysis and reliability improvement
- Engineer and enhance observability across systems and platforms
- Define, implement, and track SLIs and SLOs
- Design and build automation for recovery and selfâhealing
- Apply cloudânative resiliency and failureâisolation patterns
- Lead major incident response with an engineeringâdriven approach
- Drive systemâlevel root cause fixes
- Reduce longâterm incident volume through reliability engineering initiatives
Desired Skills
- Analyze, optimize, and enable CI/CD pipelines to improve reliability outcomes
Supplementary Skills (Good to Have)
Advanced use of AIOps for predictive reliability insights
Azure SRE lead engineers · Diverse Lynx India