Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
L

SRE

LTM
๐Ÿ‡บ๐Ÿ‡ธ United States
On-site
Staff / Principal
1 week ago
  • Incident Management
  • MuleSoft
  • Dynatrace
  • Splunk
  • Prometheus
  • Grafana
  • CI/CD
  • Jenkins
  • Git
  • Ansible
  • Python
  • PowerShell
  • Windows
  • Microservices
  • Kubernetes
  • IaC
  • Terraform
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Description

We are seeking an experienced Site Reliability Engineer SRE Lead to drive platform reliability observability and operational excellence across the API Services ecosystem

This role combines

Production engineering and reliability leadership

Platform security and vulnerability remediation

Ownership of largescale distributed runtime environments

Key responsibilities include

  • Leading reliability engineering for highscale API platforms 40K runtimes
  • Driving EOL remediation and platform stabilization efforts
  • Implementing SRE best practices
  • SLIs SLOs error budgets
  • Incident management and postmortem culture
  • Enhancing observability monitoring and proactive fault detection
  • Building resilient platforms capable of handling AIdriven usage patterns and threat models
  • Supporting global production environments with oncall and escalation coverage

Required Skills

  • Strong experience in Site Reliability Engineering Production Engineering
  • Handson expertise with
  • MuleSoft TIBCO or similar middleware platforms
  • Largescale distributed systems and runtime management
  • Deep understanding of
  • System reliability scalability and high availability design
  • Incident management root cause analysis and problem management
  • Experience with
  • Observability tools eg Dynatrace Splunk Prometheus Grafana
  • CICD pipelines Jenkins Git Ansible
  • Strong scriptingautomation skills
  • Shell Python PowerShell
  • Experience managing LinuxUnix and Windows production environments
  • Knowledge of
  • Microservices API platforms and cloudbased architectures
  • Understanding of
  • Platform security vulnerability remediation and risk mitigation in production systems
  • Excellent troubleshooting skills in highpressure realtime environments

Desired Skills

  • Experience implementing SRE frameworks SLIs SLOs error budgets
  • Familiarity with
  • Kubernetes containerized platforms
  • Infrastructure as Code Terraform Ansible
  • Exposure to
  • AIdriven operational monitoring or security tooling
  • Largescale platform modernization or migration programs
  • Middleware certifications MuleSoft or equivalent
  • Experience in regulated environments eg financial services