Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Site Reliability Engineer

EPAM Systems
  • ๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico
  • Remote
  • 1 day ago
  • Kubernetes
  • IaC
  • CI/CD
  • Incident Response
  • Azure DevOps
  • Azure Pipelines
  • Terraform
  • Ansible
  • Azure
  • Devops
  • Argo
  • NoSQL
  • Grafana
  • Elastic
  • Elastic Stack
  • Python
  • Angular
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for aSite Reliability Engineerto strengthen reliability, observability, and platform operations across cloud and Kubernetes environments. In this role, you will improve service health through automation, infrastructure as code and CI/CD practices. Apply now to help keep critical systems stable and scalable!

Responsibilities

  • Maintain service reliability by handling L2 operations and incident response
  • Operate and administer Kubernetes clusters to ensure stability and performance
  • Build and improve CI/CD workflows using Azure DevOps and Azure Pipelines
  • Automate operational tasks using scripting to reduce manual effort
  • Define and maintain infrastructure as code using Terraform and Ansible
  • Implement and refine observability using MELT signals to detect and resolve issues faster
  • Coordinate problem resolution by analyzing root causes and proposing corrective actions
  • Support secure and resilient cloud operations on Microsoft Azure
  • Document operational procedures and share knowledge to improve support readiness

Requirements

  • 2+ years of site reliability engineering or DevOps experience
  • Kubernetes administration experience supporting production workloads
  • Azure DevOps and Azure Pipelines experience delivering CI/CD workflows
  • Infrastructure as Code expertise with Terraform and Ansible
  • Proficiency in scripting languages for automation tasks
  • Strong troubleshooting skills across metrics, events, logs, and traces (MELT)
  • Strong understanding of observability concepts and tools
  • Azure fundamentals knowledge with AZ-900 or AZ-104 certification (or higher)
  • Good communication skills for cross-team incident coordination
  • English proficiency: B1+ level or higher

Nice to have

  • Argo CD administration or implementation experience
  • Experience with Apache Cassandra cluster or timeseries/NoSQL operations
  • Familiarity with Grafana and Elastic Cloud or Elastic Stack
  • HashiCorp Vault experience
  • Knowledge of programming languages, such as Python, Angular, or Go

Site Reliability Engineer ยท EPAM Systems

Auto apply with Likeremote