Site Reliability Engineer (SRE)
- Incident Response
- CI/CD
- Devops
- Python
- Java
- AWS
- Azure
- GCP
- Linux
- Unix
- IaC
- Terraform
- CloudFormation
- Docker
- Kubernetes
- Prometheus
- Grafana
- Datadog
- Incident Management
Site Reliability Engineer (SRE)
Hartford, CT
Job Summary
We are seeking a highly skilledSite Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our systems and services. The SRE will bridge development and operations by applying software engineering principles to infrastructure and operations problems.
Key Responsibilities
-
Design, implement, and maintain highly available and scalable systems.
-
Monitor system performance and troubleshoot issues across production environments.
-
Automate operational tasks and infrastructure provisioning.
-
Improve system reliability through proactive testing and observability.
-
Collaborate with development teams to enhance application performance and resilience.
-
Lead incident response, root cause analysis (RCA), and postmortems.
-
Define and monitor SLIs, SLOs, and SLAs.
-
Manage CI/CD pipelines and deployment processes.
-
Implement security best practices across infrastructure and services.
-
Reduce toil by automating repetitive operational tasks.
Required Qualifications
-
Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).
-
3+ years of experience in SRE, DevOps, or system engineering roles.
-
Strong programming skills (Python, Go, Java, or similar).
-
Experience with cloud platforms (AWS, Azure, or GCP).
-
Proficiency in Linux/Unix systems administration.
-
Experience with Infrastructure as Code (Terraform, CloudFormation).
-
Familiarity with containerization and orchestration (Docker, Kubernetes).
-
Experience with monitoring tools (Prometheus, Grafana, Datadog, etc.).
-
Strong understanding of networking fundamentals.
-
Excellent problem-solving and communication skills.
Preferred Qualifications
-
Experience managing large-scale distributed systems.
-
Knowledge of security best practices and compliance standards.
-
Experience with incident management frameworks.
-
Certifications in cloud technologies (AWS, GCP, Azure)
Site Reliability Engineer (SRE) ยท Info Way Solutions LLC