SRE- India
- Incident Management
- GCP
- GKE
- Kubernetes
- Docker
- Terraform
- Ansible
- Prometheus
- Grafana
- Datadog
- Splunk
- OpenTelemetry
- Python
TheSite Reliability Engineer (SRE) will be responsible for ensuring the reliability, availability, performance, and scalability of business-critical applications and cloud infrastructure. The role will work closely with Development, Cloud, Platform, and Operations teams to improve system resilience, strengthen observability practices, and drive operational excellence across production environments. The engineer will support incident management activities, perform root cause analysis, implement preventive actions, and contribute to continuous service improvement initiatives.
The role will focus on monitoring and maintaining production systems, defining and tracking Service Level Indicators (SLIs) and Service Level Objectives (SLOs), and ensuring high levels of system uptime and performance. The engineer will leverage cloud-native technologies and automation frameworks to reduce manual operational effort, enhance deployment reliability, and improve overall platform stability. Responsibilities include infrastructure automation, capacity planning, performance optimization, observability enhancement, and proactive risk identification.
The ideal candidate should possess hands-on experience with Google Cloud Platform (GCP), GKE/Kubernetes, Docker, Terraform, Ansible, and observability tools such as Prometheus, Grafana, Datadog, Splunk, or OpenTelemetry. Strong scripting and automation skills using Python, Go, or Shell are essential. The role requires excellent analytical and troubleshooting abilities, along with strong collaboration and communication skills to partner effectively with multiple technology teams and ensure reliable delivery of business services in a fast-paced production environment.
SRE- India · Diverse Lynx India