DL
SRE
Diverse Lynx India
🇮🇳 India
Hybrid
1 week ago
- Incident Management
- GCP
- GKE
- Kubernetes
- Docker
- Terraform
- Ansible
- Prometheus
- Grafana
- Datadog
- Splunk
- OpenTelemetry
- Python
- Disaster Recovery
1 week ago
| Description: TheSite Reliability Engineer (SRE) will be responsible for ensuring the reliability, availability, performance, and scalability of business-critical applications and cloud infrastructure. The role will work closely with Development, Cloud, Platform, and Operations teams to improve system resilience, strengthen observability practices, and drive operational excellence across production environments. The engineer will support incident management activities, perform root cause analysis, implement preventive actions, and contribute to continuous service improvement initiatives. The role will focus on monitoring and maintaining production systems, defining and tracking Service Level Indicators (SLIs) and Service Level Objectives (SLOs), and ensuring high levels of system uptime and performance. The engineer will leverage cloud-native technologies and automation frameworks to reduce manual operational effort, enhance deployment reliability, and improve overall platform stability. Responsibilities include infrastructure automation, capacity planning, performance optimization, observability enhancement, and proactive risk identification. The ideal candidate should possess hands-on experience with Google Cloud Platform (GCP), GKE/Kubernetes, Docker, Terraform, Ansible, and observability tools such as Prometheus, Grafana, Datadog, Splunk, or OpenTelemetry. Strong scripting and automation skills using Python, Go, or Shell are essential. The role requires excellent analytical and troubleshooting abilities, along with strong collaboration and communication skills to partner effectively with multiple technology teams and ensure reliable delivery of business services in a fast-paced production environment. |
|||
|
MAH | PUNE | ||
|
8 + Years | ||
|
• Google Cloud Platform (GCP) & Kubernetes (GKE) - Strong hands-on experience in managing, troubleshooting, and optimizing cloud-native applications and containerized workloads. • Observability & Reliability Engineering - Expertise in monitoring, alerting, incident management, root cause analysis, and tools such as Prometheus, Grafana, Datadog, Splunk, and OpenTelemetry. • Infrastructure Automation & Scripting - Hands-on experience with Terraform, Ansible, and automation using Python, Go, or Shell scripting to improve operational efficiency and platform reliability. |
||
|
Multi-Cloud & Cloud Security Knowledge – Experience with cloud governance, security best practices, disaster recovery, capacity planning, and resilience patterns across large-scale cloud environments. • OpenTelemetry & Advanced Observability – Hands-on experience implementing distributed tracing, synthetic monitoring, and observability frameworks using OpenTelemetry and modern monitoring platforms |
||
|
Reatil | ||
|
8 + Years | ||
|
Video Interview | ||
|
Hybrid |
SRE · Diverse Lynx India