Site Reliability Engineer (SRE)
- CI/CD
- Devops
- Terraform
- Ansible
- Jenkins
- GitHub Actions
- Kubernetes
- Docker
- AWS
- Azure
- GCP
- Bash
- Python
- Node.js
- Load Balancing
- Prometheus
- Grafana
- Datadog
- Splunk
The ideal candidate will have strong experience with cloud technologies, automation frameworks, CI/CD pipelines, and monitoring systems, and will collaborate across development, DevOps, and architecture teams to embed reliability into every phase of delivery.
Key Responsibilities
-
Design and build automated deployment systems and infrastructure that enhance reliability and consistency across cloud environments
-
Develop self-service tools and frameworks to streamline infrastructure and application management tasks
-
Collaborate with software engineering and DevOps teams to ensure highly available and resilient systems
-
Implement automation using Terraform, Ansible, or similar tools to enable repeatable and scalable operations
-
Manage and optimize CI/CD pipelines using tools such as Jenkins, GitHub Actions, or Codefresh
-
Configure, monitor, and maintain Kubernetes, Docker, and cloud-based services (AWS, Azure, GCP)
-
Integrate observability and monitoring solutions to proactively detect and resolve reliability issues
-
Conduct root cause analysis and implement corrective measures to prevent recurrence
-
Create and maintain documentation for infrastructure, systems, and operational processes
-
Continuously evaluate emerging technologies and recommend improvements to platform reliability and performance
Required Qualifications
-
Bachelor's degree in Computer Science, Engineering, or related field — Master's preferred
-
8+ years of experience in infrastructure, DevOps, or platform engineering roles
-
Hands-on experience with AWS, Azure, or GCP cloud environments
-
Strong proficiency in CI/CD pipelines and automation tools (Jenkins, GitHub Actions, Codefresh, or equivalent)
-
Expertise in Terraform, Kubernetes, and Docker
-
Scripting experience in Bash, with additional experience in Python or Node.js preferred
-
Deep understanding of high availability, load balancing, clustering, and reliability engineering concepts
-
Experience with monitoring and observability tools (Prometheus, Grafana, Datadog, or Splunk)
-
Strong problem-solving, analytical, and troubleshooting skills
-
Excellent written and verbal communication skills, with the ability to collaborate across technical and leadership teams
Preferred Qualifications
-
Cloud certifications (AWS, Azure, or GCP)
-
Experience implementing chaos engineering or resiliency testing practices
-
Familiarity with container orchestration at scale and hybrid cloud architectures
Compensation
Competitive compensation and comprehensive benefits
Location
Philadelphia, PA
Site Reliability Engineer (SRE) · CrossTech Consulting Group, Inc