Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
IW

Lead Site Reliability Engineer – Observability

Info Way Solutions LLC
Location not stated
Staff / Principal
2 months ago
  • Splunk
  • Elasticsearch
  • Grafana
  • OpenTelemetry
  • Prometheus
  • Kibana
  • Terraform
  • IaC
  • Devops
  • Elastic Stack
  • Python
  • Ruby
  • Bash
  • Kubernetes
  • Docker
  • AWS
  • Azure
  • GCP
  • Ansible
  • CI/CD
  • FedRAMP
  • Linux
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Title: Lead Site Reliability Engineer – Observability

Location: Remote (United States)
Job Type: Contract

Job Summary

We are seeking a Lead Site Reliability Engineer (SRE) – Observability to join our growing platform engineering team. The ideal candidate will have extensive experience designing, implementing, and operating enterprise observability platforms across large-scale cloud environments. This role will focus on logging, metrics, distributed tracing, and alerting solutions while driving reliability, scalability, and operational excellence.


Key Responsibilities

  • Design, deploy, and manage enterprise observability platforms.
  • Administer and support Splunk Enterprise and Splunk Cloud environments, including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and maintain large-scale Elasticsearch (ELK) clusters for log analytics and search.
  • Implement distributed tracing solutions using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and retention strategies.
  • Develop and optimize monitoring solutions leveraging Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry.
  • Create dashboards, alerts, analytics, and visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure provisioning and management using Terraform and Infrastructure as Code (IaC).
  • Collaborate with engineering, operations, and security teams to improve platform reliability and performance.
  • Drive best practices for observability, automation, and operational excellence.

Required Skills

  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience with Splunk Enterprise and/or Splunk Cloud administration.
  • Strong expertise in Splunk SPL.
  • Experience with Elasticsearch, ELK Stack, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, and Kafka.
  • Strong understanding of modern observability practices, including metrics, logs, and distributed tracing.
  • Experience with Terraform and Infrastructure as Code (IaC).
  • Programming experience in Python, Go, Ruby, or Bash.
  • Strong troubleshooting, communication, and leadership skills.

Preferred Skills

  • Splunk Certification.
  • Experience with Kubernetes and Docker.
  • Exposure to AWS, Azure, or GCP cloud platforms.
  • Experience with Ansible, Consul, and CI/CD pipelines.
  • Familiarity with service mesh technologies.
  • Experience supporting FedRAMP High or IL-5 regulated environments.

Technology Stack

  • Splunk Enterprise
  • Splunk Cloud
  • Elasticsearch (ELK Stack)
  • Kibana
  • Prometheus
  • Grafana
  • Grafana Tempo
  • OpenTelemetry
  • Kafka
  • Terraform
  • Kubernetes
  • Docker
  • Linux
  • Python, Go, Ruby, Bash
  • AWS, Azure, GCP
  • Ansible
  • Consul

Work Authorization Requirements

  • Must be a U.S. Citizen or U.S. National.
  • Work must be performed within the United States.
  • Candidates must be eligible to support FedRAMP High and IL-5 environments.

Experience

  • 7+ years of relevant experience in SRE, DevOps, or Platform Engineering.
  • Experience leading observability initiatives in enterprise-scale environments is highly preferred.

This is an excellent opportunity to lead critical observability initiatives and build scalable, resilient platforms supporting mission-critical cloud infrastructure across the United States.

Lead Site Reliability Engineer – Observability · Info Way Solutions LLC

Auto apply with Likeremote