SRE
- AWS
- Dynatrace
- Prometheus
- CloudWatch
- ServiceNow
- Jira
- GitLab
- Jenkins
- Packer
- CI/CD
- Incident Management
- Kubernetes
- Docker
- Python
- Java
- Node.js
- Terraform
- AI
- Machine Learning
- RAG
- AI/ML
- VPC
- EC2
- IAM
- IaC
- System Design
|
Site Reliability Engineer (SRE) – 5 to 9 Years Experience Experience: 5–9 years in Site Reliability Engineering. Skills Required: Strong experience in infrastructure management, patching, monitoring, troubleshooting, and cloud operations. Hands-on knowledge of AWS, Dynatrace, Prometheus, CloudWatch, ServiceNow, and JIRA. Experience with GitLab, Harbor, Jenkins, XL Deploy, Camunda, Packer, Rancher, CI/CD pipelines, and incident management processes. Basic understanding of Kubernetes and Docker. Proficiency in automation using Python, Java, or NodeJS; Terraform and AWS certification preferred. Good understanding of SRE fundamentals (SLI/SLO/SLA), AI/GenAI concepts, Agentic AI, LLMs, RAG pipelines, vector databases, and AI platform operations. Strong communication skills with a proactive approach to performance optimization and problem-solving. Responsibilities: Monitor and maintain cloud, infrastructure, application, and AI platforms. Support AI/ML workloads, LLM applications, and production services by ensuring reliability, availability, performance, and operational readiness. Manage infrastructure provisioning, upgrades, patching, capacity planning, and resource optimization (CPU, memory, storage, GPU, APIs). Automate operational tasks, support deployments, create symptom-based monitoring and alerting, and drive continuous improvements. Respond to incidents, participate in on-call support, manage critical incident bridges, conduct post-incident reviews, and collaborate with development teams to improve testing, releases, scalability, and service reliability. Senior Site Reliability Engineer (Lead SRE) – 10+ Years Experience Experience: Minimum 10+ years in SRE, Cloud, Platform Engineering, or Production Operations. Skills Required: Strong expertise in infrastructure management, cloud operations, monitoring, incident management, and platform reliability. Hands-on experience with Dynatrace, Prometheus, CloudWatch, JIRA, ServiceNow, GitLab, Harbor, Jenkins, XL Deploy, Camunda, Packer, and Rancher. Good knowledge of AWS services (VPC, ALB, EC2, S3, IAM, CloudWatch), Kubernetes, Docker, CI/CD, Terraform, and automation using Python/Java/NodeJS. Experience building scalable, highly available, fault-tolerant platforms and implementing Infrastructure as Code. Strong understanding of eCommerce platform performance and stability challenges. Excellent communication skills with a proactive approach to problem-solving, performance optimization, cost management, and platform modernization. AWS certification is preferred. Responsibilities: Lead SRE and observability strategy with Dynatrace, ensuring end-to-end visibility across applications, infrastructure, and business journeys. Drive automation, self-healing solutions, runbooks-as-code, auto-remediation, and reusable Terraform modules. Manage critical incidents, bridge calls, post-incident reviews, and reliability governance. Define SRE standards, KPIs, SLI/SLO/SLA targets, and operating models. Partner with development teams on system design, testing, releases, scalability, and capacity planning. Build proactive monitoring and intelligent alerting using anomaly detection and SLO-based thresholds. Align reliability metrics with business KPIs, conduct root cause analysis, engage with senior stakeholders, and provide executive-level updates to client and leadership teams. |
SRE · Diverse Lynx India