Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead Site Reliability Engineer

EPAM Systems
  • ๐Ÿ‡ง๐Ÿ‡ท Brazil | ๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ด Colombia
  • Remote
  • Staff / Principal
  • 14 hours ago
  • Devops
  • CI/CD
  • GitLab
  • Python
  • Kubernetes
  • IAM
  • Incident Response
  • AWS
  • Azure
  • AI
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for aLead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate safe delivery of change in a fast-moving environment. You will elevate DevOps maturity across tooling, processes, and engineering practices while solving complex reliability challenges.

Responsibilities

  • Own reliability outcomes for critical infrastructure and set SRE engineering standards
  • Design and implement scalable DevOps tooling and processes that improve delivery speed and safety
  • Build and maintain CI/CD workflows and source control practices using GitLab where appropriate
  • Automate infrastructure operations with Python to reduce toil and improve consistency
  • Improve Kubernetes-based workflows to support dependable deployments and runtime stability
  • Diagnose and resolve business-critical incidents during on-call shifts with urgency and rigor
  • Evaluate and harden infrastructure domains such as networking, compute, security, IAM, and configuration automation
  • Drive enterprise-grade release management practices across environments
  • Partner with stakeholders to translate reliability needs into actionable engineering plans
  • Promote long-term engineering solutions over quick fixes through root-cause analysis and follow-up actions

Requirements

  • 5+ years of site reliability engineering experience in production environments
  • 5+ years of cloud platform experience with a leading provider
  • Proven leadership skills to drive reliability improvements and mentor engineers
  • Enterprise-scale release management experience across complex systems
  • Strong CI/CD knowledge covering pipelines, source control, and automation practices
  • Advanced Python programming skills for automation and engineering solutions
  • Hands-on Kubernetes experience as a developer in delivery workflows
  • Strong infrastructure fundamentals across networking, compute, security, and IAM
  • Excellent analytical skills for strategic thinking and complex problem solving
  • Effective incident response skills, including on-call ownership for critical issues
  • English proficiency: B2 Upper-Intermediate

Nice to have

  • Amazon Web Services expertise in production environments
  • Microsoft Azure expertise in production environments
  • AI Architecture experience for reliability-aware AI platform design
  • AI Solution Engineering experience supporting clients with scalable implementations
  • Gen AI Solutions Development experience in enterprise settings

Lead Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote