Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Site Reliability Engineer

EPAM Systems
๐Ÿ‡ต๐Ÿ‡ฑ Poland
Hybrid
Senior
1 month ago
  • AI
  • IaC
  • CI/CD
  • Azure
  • Incident Management
  • Terraform
  • Bicep
  • Python
  • Bash
  • FinOps
  • triage
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aSenior Site Reliability Engineerto own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.

Responsibilities

  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect

Requirements

  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
  • Knowledge of FinOps basics for AI workloads
  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code

Senior Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote