A
SRE
Apolis
πΊπΈ United States
Hybrid
1 month ago
$42 β $44 / hour
- Devops
- Python
- React.js
- Java
- Prometheus
- Grafana
- OpenTelemetry
- Loki
- Splunk
- Elasticsearch
- Node.js
- GCP
- Kubernetes
- AI
- Apache Airflow
- SQL
- PostgreSQL
- triage
- Istio
- Envoy
- IaC
- Terraform
- Ansible
- Incident Response
- System Design
- Incident Management
1 month ago
We are looking for a strong SRE resource to work from Rhode Island office 3 days a week (Hybrid). Share strong profile aligning to the JD listed below.
Location : Rhode Island
Required Qualifications
- 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
- Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents β structured leadership updates, not just participant involvement
- Experience tuning and validating time-series anomaly detection models in a production observability context β this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
- Strong programming proficiency in Python , React , and Java at production quality β capable of writing operational tooling that other engineers will rely on
- Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services
- Deep observability platform experience: Prometheus , Grafana , OpenTelemetry , and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
- Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
- Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
- Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.
- Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal .
Preferred Qualifications
- Experience owning Production Readiness Reviews or service launch gates.
- Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.
- Hands-on chaos or fault injection experience.
- TIC (Technical Incident Commander) certification or equivalent structured incident command training
- Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
- LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) β design or implementation experience
- Experience with streaming data platforms: Kafka .
- Experience with service mesh and traffic management: Istio , Envoy .
- Infrastructure-as-code proficiency at production scale: Terraform or Ansible
Position Summary :
- Own and drive the end-to-end reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and on-premises environments.
- Establish and maintain SLI/SLO health, alerting strategies, observability standards, and business-aligned monitoring for the assigned application domain.
- Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
- Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
- Build and optimize automation, self-service capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
- Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
- Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement.
- Influence organizational adoption of SLO-driven engineering, observability, incident management, and reliability practices through collaboration, credibility, and measurable outcomes.
SRE Β· Apolis