DL
SRE DEVOPS
Diverse Lynx India
🇮🇳 India
On-site
2 months ago
- Incident Management
- AWS
- Azure
- IaC
- Microservices
- Docker
- Kubernetes
- GCP
- Terraform
- CI/CD
- Git
- Python
- Bash
2 months ago
Experience: 6-8 years
Must be ready to work from office 5 days a week.
Responsibilities:
- Define, track, and report on SLOs and SLIs for critical service
- Setup Monitoring and observability for the system​
- Take lead on complex incidents and provide deep technical expertise to resolve issues quickly.​
- Perform RCA in-depth for incident management and suggest permanent fix
- Participate in design reviews focusing on reliability and scalability​
- Design and implement automation for high-availability systems and fault-tolerant architectures​
- Design, build, and manage scalable cloud infrastructure (AWS/Azure) via IAC
- Deploy and orchestrate microservices using Docker and Kubernetes.
Skills:
- Expertise inSite Reliability Engineering (SRE)processes and skills.
- Solid understanding of Service Level Agreements (SLA), Service Level Indicators (SLI), and Service Level Objectives (SLO).
- Extensive experience withinfrastructure monitoring, Application Performance Management (APM), and observability tools.
- Proficient inIncident and Problem Management.
- Proven track record in contributing to toil reduction.
- Strongtroubleshooting skills, including Root Cause Analysis (RCA) and postmortem processes.
- Expertise in any one of the cloud providers (AWS/Azure/GCP), Kubernetes, Terraform, CICD, Git and Scripting (Python/Bash)
SRE DEVOPS · Diverse Lynx India