Lead II - DevOps Engineering SRE DevOps
- Devops
- AWS
- Kubernetes
- CI/CD
- GitLab
- triage
- Splunk
- Terraform
- GitLab CI/CD
- IaC
- Incident Management
- Linux
- Unix
- EKS
- Helm
- Prometheus
- Grafana
- CloudWatch
- Docker
- Python
- Bash
- Microservices
- IAM
Experience: 8-15 Years
Job Location: Chennai, Bangalore, Hyderabad, Kochi, Trivandrum, Noida, Pune
DevOps SRE Engineer / Platform EngineerRole Overview
We are looking for highly hands-onDevOps SRE Engineerswith strongPlatform Engineeringexperience supporting application and API hosting platforms inAWS and Kubernetesenvironments.
The ideal candidate should have strong recent hands-on technical expertise in production environments, with the ability to operate, troubleshoot, optimize, and enhance platforms within amulti-vendor ecosysteminvolving multiple teams.
This is ahands-on individual contributor role and not a people-management position. Strong communication skills are essential, along with the ability to lead technical discussions with Tier 2 teams, stakeholders, and cross-functional engineering groups.
Key Responsibilities
- Maintain, troubleshoot, and enhanceCI/CD pipelines, withGitLab preferred.
- Manage and supportAWS infrastructureand cloud-native services.
- Deploy, operate, and troubleshootKubernetes workloads and platforms.
- Performproduction incident triage, debugging, and root cause analysis.
- Uselogs, metrics, monitoring tools, and Splunkfor effective production troubleshooting.
- Develop and enhanceTerraform-based infrastructure and automation.
- Supportrepository restructuring, platform modernization, and engineering transformation initiatives.
- Drive improvements inplatform stability, reliability, observability, scalability, and performance.
- Support platforms that host and enableapplication and API services.
- Collaborate closely with theGCC team and UST BFF leads.
- Work effectively across a multi-vendor engineering ecosystem and coordinate technical resolution across teams.
- Drive technical discussions, identify recurring platform issues, and implement sustainable solutions.
- Contribute to continuous improvement of platform engineering practices, automation, and operational processes.
Must-Have Skills
- Strong hands-on experience inDevOps / Site Reliability Engineering (SRE).
- StrongAWSexperience, including production infrastructure management and troubleshooting.
- Strong hands-onKubernetesexperience.
- Strong experience withCI/CD pipelines, preferablyGitLab CI/CD.
- Hands-on experience withTerraformand Infrastructure as Code (IaC).
- Strong production troubleshooting andincident managementexperience.
- Experience withSplunk, logs, metrics, and monitoring/observability tools.
- Strong understanding ofLinux/Unixenvironments and troubleshooting.
- Experience supportingapplication and API hosting platforms.
- Experience with platform reliability, availability, performance, and observability improvements.
- Strong debugging androot cause analysis (RCA)capabilities.
- Ability to work independently as a hands-on technical contributor.
- Strong communication and stakeholder-management skills.
- Ability to drive technical discussions withTier 2 support, engineering teams, stakeholders, and cross-functional teams.
- Experience working in amulti-vendor / distributed engineering environment.
Good-to-Have Skills
- Experience withAWS EKSor other managed Kubernetes services.
- Experience withGitLab administration or advanced GitLab CI/CD.
- Advanced Terraform modules, state management, and automation experience.
- Experience withHelmand Kubernetes deployment automation.
- Experience with observability platforms such asPrometheus, Grafana, CloudWatch, or similar tools.
- Experience withDocker/containerization.
- Experience with scripting/automation usingPython, Shell, or Bash.
- Knowledge ofAPI platforms, microservices, and cloud-native architectures.
- Experience with platform modernization and repository restructuring.
- Experience implementingSRE practices, SLIs/SLOs, error budgets, and reliability engineering principles.
- Experience with security, IAM, networking, and cost optimization in AWS.
- Experience coordinating technical initiatives across multiple vendors and engineering teams.
Experience Range
6–12 yearsof overall experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or related roles.
Preferred:4+ years of strong hands-on experience with AWS and Kubernetes in production environments.
Candidates with extensive recent hands-on expertise and strong production troubleshooting experience will be preferred over candidates with primarily managerial or coordination experience.
Preferred Candidate Profile
The ideal candidate is a strong hands-on engineer who canoperate production platforms, troubleshoot complex incidents, automate infrastructure, improve reliability, and work across multiple engineering/vendor teams. The role requires someone who can not only execute technically but also confidently drive technical conversations and influence stakeholders toward effective platform solutions.
Lead II - DevOps Engineering SRE DevOps · UST