Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
A

Site Reliability Engineer (SRE) – Production Services

Apolis
πŸ‡ΊπŸ‡Έ United States
On-site
Mid level
2 months ago
$50 – $60 / hour
  • Java
  • Spring Boot
  • Apache Kafka
  • Devops
  • CI/CD
  • AI
  • Jenkins
  • GitLab CI
  • Azure DevOps
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Title:Site Reliability Engineer (SRE) – Production Services

Location:Pittsburgh, PA 15219 (Onsite)/Local Candidates Only

Tax Term (W2, C2C):W2

Job Type (Permanent/Contract):Contract

Duration:Long Term

Description:

We are seeking an experiencedSite Reliability Engineer (SRE) – Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise inJava Spring Boot, Apache Kafka, DevOps, and CI/CD automation, along with experience building scalable, resilient, and highly available production systems. This is a fully onsite role inPittsburgh, PA, and only local candidates will be considered.

Role and Responsibilities:

  • Automate high-volume production support requests and operational workflows.
  • Develop self-service and agent-driven automation solutions to minimize manual effort.
  • Implement standardized operational processes with auditability and resilience.
  • Build auto-retry, backoff, and recovery mechanisms for recurring production failures.
  • Define, monitor, and maintain Service Level Objectives (SLOs) and apply error budget principles.
  • Improve reliability of batch processing through standardized recovery patterns.
  • Develop observability dashboards for incidents, failures, automation coverage, and operational metrics.
  • Create and enhance production runbooks and convert them into automated remediation workflows.
  • Drive permanent resolution of recurring production issues through root cause analysis.
  • Implement self-healing capabilities to reduce operational intervention.
  • Optimize monitoring and alerting platforms (Moogsoft or similar) to improve signal-to-noise ratio.
  • Leverage automation and AI-driven operational solutions for recurring production issues.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and production stability.

Required Skills:

  • 12+ years of overall IT experience.
  • 8–10+ years of Site Reliability Engineering (SRE) or Production Support experience.
  • Strong hands-on experience withJava andSpring Boot.
  • Experience withApache Kafka.
  • Strong knowledge ofDevOps practices and tools.
  • Expertise inCI/CD automation (Jenkins, GitLab CI, Azure DevOps, etc.).
  • Experience with production monitoring, observability, dashboards, and alerting tools.
  • Knowledge of Service Level Objectives (SLOs), SLIs, and Error Budgets.
  • Experience implementing automation, self-healing, and operational runbooks.
  • Strong troubleshooting and root cause analysis skills.
  • Experience working in enterprise production support environments.

Qualifications:

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • Experience with cloud platforms and container technologies is a plus.
  • Excellent communication and collaboration skills.
  • Ability to work in a fast-paced production support environment.
  • Local candidates available to work onsite in Pittsburgh, PA.

Site Reliability Engineer (SRE) – Production Services Β· Apolis

Auto apply with Likeremote