Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
L

SRE Engineer

LTM
🇺🇸 United States
On-site
1 month ago
  • AWS
  • EC2
  • VPC
  • RDS
  • EKS
  • Incident Management
  • CloudWatch
  • Dynatrace
  • CI/CD
  • Unix
  • Linux
  • triage
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

#Careers

JC1483593

Qualifications

·       Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.

·       Hands-on experience with incident management and 24/7 production support models.

·       Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.

·       Experience building and maintaining monitoring dashboards.

·       Strong troubleshooting skills across infrastructure, networking, and application layers.

·       Working knowledge of CI/CD pipelines and AWS deployment processes.

·       Experience working with databases and Unix/Linux environments.

Key Responsibilities

Incident Management and Production Support

·       Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.

·       Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.

·       Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.

·       Participate in on-call rotations, major incident bridges, and post-incident reviews.

·       Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.

Monitoring and Operational Health

·       Perform regular health checks across applications, infrastructure, and AWS services.

·       Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes.

·       Respond proactively to s related to resource utilization, latency, errors, and availability.

·       Maintain and improve monitoring and observability dashboards.

SRE Engineer · LTM

Auto apply with Likeremote