ST
Sr. SRE engineer
Spark Tek Inc
🇺🇸 United States
On-site
Senior
1 month ago
- AWS
- EC2
- VPC
- RDS
- EKS
- Incident Management
- CloudWatch
- Dynatrace
- CI/CD
- Unix
- Linux
- triage
1 month ago
Location: Atlanta, GA       
Qualifications
•     Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.
•     Hands-on experience with incident management and 24/7 production support models.
•     Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.
•     Experience building and maintaining monitoring dashboards.
•     Strong troubleshooting skills across infrastructure, networking, and application layers.
•     Working knowledge of CI/CD pipelines and AWS deployment processes.
•     Experience working with databases and Unix/Linux environments.
Key Responsibilities
Incident Management and Production Support
•     Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.
•     Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
•     Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
•     Participate in on-call rotations, major incident bridges, and post-incident reviews.
•     Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.
Monitoring and Operational Health
•     Perform regular health checks across applications, infrastructure, and AWS services.
•     Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and ThousandEyes.
•     Respond proactively to alerts related to resource utilization, latency, errors, and availability.
•     Maintain and improve monitoring and observability dashboards."               
Sr. SRE engineer · Spark Tek Inc