A
Site Reliability Engineer (SRE) β Production Services
Apolis
πΊπΈ United States
On-site
Mid level
2 months ago
$50 β $60 / hour
- Java
- Spring Boot
- Apache Kafka
- Devops
- CI/CD
- AI
- Jenkins
- GitLab CI
- Azure DevOps
2 months ago
Job Title:Site Reliability Engineer (SRE) β Production Services
Location:Pittsburgh, PA 15219 (Onsite)/Local Candidates Only
Tax Term (W2, C2C):W2
Job Type (Permanent/Contract):Contract
Duration:Long Term
Description:
We are seeking an experiencedSite Reliability Engineer (SRE) β Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise inJava Spring Boot, Apache Kafka, DevOps, and CI/CD automation, along with experience building scalable, resilient, and highly available production systems. This is a fully onsite role inPittsburgh, PA, and only local candidates will be considered.
Role and Responsibilities:
- Automate high-volume production support requests and operational workflows.
- Develop self-service and agent-driven automation solutions to minimize manual effort.
- Implement standardized operational processes with auditability and resilience.
- Build auto-retry, backoff, and recovery mechanisms for recurring production failures.
- Define, monitor, and maintain Service Level Objectives (SLOs) and apply error budget principles.
- Improve reliability of batch processing through standardized recovery patterns.
- Develop observability dashboards for incidents, failures, automation coverage, and operational metrics.
- Create and enhance production runbooks and convert them into automated remediation workflows.
- Drive permanent resolution of recurring production issues through root cause analysis.
- Implement self-healing capabilities to reduce operational intervention.
- Optimize monitoring and alerting platforms (Moogsoft or similar) to improve signal-to-noise ratio.
- Leverage automation and AI-driven operational solutions for recurring production issues.
- Collaborate with development, infrastructure, and operations teams to improve system reliability and production stability.
Required Skills:
- 12+ years of overall IT experience.
- 8β10+ years of Site Reliability Engineering (SRE) or Production Support experience.
- Strong hands-on experience withJava andSpring Boot.
- Experience withApache Kafka.
- Strong knowledge ofDevOps practices and tools.
- Expertise inCI/CD automation (Jenkins, GitLab CI, Azure DevOps, etc.).
- Experience with production monitoring, observability, dashboards, and alerting tools.
- Knowledge of Service Level Objectives (SLOs), SLIs, and Error Budgets.
- Experience implementing automation, self-healing, and operational runbooks.
- Strong troubleshooting and root cause analysis skills.
- Experience working in enterprise production support environments.
Qualifications:
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- Experience with cloud platforms and container technologies is a plus.
- Excellent communication and collaboration skills.
- Ability to work in a fast-paced production support environment.
- Local candidates available to work onsite in Pittsburgh, PA.
Site Reliability Engineer (SRE) β Production Services Β· Apolis