DL
SRE/Triaging Engineer
Diverse Lynx India
Location not stated
2 months ago
- AWS
- Microservices
- Incident Response
- triage
- ECS
- Fargate
- API Gateway
- SNS
- SQS
- PostgreSQL
- Datadog
- CloudWatch
- GitHub Actions
- Terraform
- JMeter
- ITIL
- Incident Management
2 months ago
Essential Skills: L2 - SRE/Triaging Engineer (AWS | Microservices | Incident Response)First technical responder for high impact production alerts| leading triage| restoring services using runbooks| and isolating faults across infrastructure| platform| and application layers. This is a hands on operational role focused on uptimenot software development.Responsibilities Real time response to P1/P2/P3 alerts with strong situational awareness. Rapid triage: validate alerts| assess impact| correlate logs/metrics/traces. Service restoration using runbooks (restarts| scaling| failover). Technical fault isolation across AWS infra| platform services| and application behaviour. Clear| structured escalation with timestamps| logs| metrics| and attempted steps. Continuous improvement of runbooks| alert hygiene| and operational processes. Participation in a 247 on call rotation.Technical EnvironmentAWS Compute & Integration ECS (Fargate)| Lambda| API Gateway| SNS/SQS| MSK (Kafka)| Amazon MQ| IBM MQ| EventBridgeDatastores PostgreSQL| DynamoDBObservability Datadog (metrics + APM)| CloudWatch Logs| SplunkCI/CD & Infra Awareness GitHub Actions| Terraform| ECR| deployment rollback workflowsMust Have Experience36 years in Production Support| Cloud Ops| or SRE (L2).Strong AWS troubleshooting skills across compute| networking| and platform services.Experience supporting microservices and event driven architectures.Strong log analysis and incident triage capability.Ability to create/improve runbooks and reduce alert noise. Nice to HavePython scripting for automation.Exposure to performance triage (e.g.| JMeter).Familiarity with ITIL or structured incident management processes.
SRE/Triaging Engineer · Diverse Lynx India