
Senior Site Reliability Engineer
RapidAI
🇮🇳 India
Hybrid
Senior
3 weeks ago
- AI
- Incident Response
- EKS
- IaC
- Terraform
- Helm
- GitOps
- Load Testing
- AWS
- Devops
- EC2
- VPC
- IAM
- RDS
- CloudWatch
- Kubernetes
- Linux
- Prometheus
- Grafana
- Jaeger
- Python
- Bash
3 weeks ago
- Own the availability, performance, and incident response for Rapid's production EKS clusters
- Design and operate the full observability stack — metrics, logs, traces — with
Open Telemetry as the foundation - Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
- Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
- Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
- Tune autoscaling, networking, and cost efficiency across AWS workloads
- On-call rotation with the expectation you'll also fix the underlying cause, not just the alert  What We Looking For:
- 10+ years in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
surrounding ecosystem - Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
- Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
- Strong foundation in Linux, networking, and distributed systems fundamentals
- Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out - Startup mindset: you make decisions with incomplete information and iterate quickly
What You Do:
Senior Site Reliability Engineer · RapidAI