Lead Site Reliability Engineer – Observability
- Splunk
- Elasticsearch
- Grafana
- OpenTelemetry
- Prometheus
- Kibana
- Terraform
- IaC
- Devops
- Elastic Stack
- Python
- Ruby
- Bash
- Kubernetes
- Docker
- AWS
- Azure
- GCP
- Ansible
- CI/CD
- FedRAMP
- Linux
Job Title: Lead Site Reliability Engineer – Observability
Location: Remote (United States)
Job Type: Contract
Job Summary
We are seeking a Lead Site Reliability Engineer (SRE) – Observability to join our growing platform engineering team. The ideal candidate will have extensive experience designing, implementing, and operating enterprise observability platforms across large-scale cloud environments. This role will focus on logging, metrics, distributed tracing, and alerting solutions while driving reliability, scalability, and operational excellence.
Key Responsibilities
- Design, deploy, and manage enterprise observability platforms.
- Administer and support Splunk Enterprise and Splunk Cloud environments, including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
- Deploy and maintain large-scale Elasticsearch (ELK) clusters for log analytics and search.
- Implement distributed tracing solutions using Grafana Tempo and OpenTelemetry.
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and retention strategies.
- Develop and optimize monitoring solutions leveraging Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry.
- Create dashboards, alerts, analytics, and visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
- Automate infrastructure provisioning and management using Terraform and Infrastructure as Code (IaC).
- Collaborate with engineering, operations, and security teams to improve platform reliability and performance.
- Drive best practices for observability, automation, and operational excellence.
Required Skills
- 7+ years of experience in Site Reliability Engineering, Platform Engineering, or DevOps.
- Hands-on experience with Splunk Enterprise and/or Splunk Cloud administration.
- Strong expertise in Splunk SPL.
- Experience with Elasticsearch, ELK Stack, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, and Kafka.
- Strong understanding of modern observability practices, including metrics, logs, and distributed tracing.
- Experience with Terraform and Infrastructure as Code (IaC).
- Programming experience in Python, Go, Ruby, or Bash.
- Strong troubleshooting, communication, and leadership skills.
Preferred Skills
- Splunk Certification.
- Experience with Kubernetes and Docker.
- Exposure to AWS, Azure, or GCP cloud platforms.
- Experience with Ansible, Consul, and CI/CD pipelines.
- Familiarity with service mesh technologies.
- Experience supporting FedRAMP High or IL-5 regulated environments.
Technology Stack
- Splunk Enterprise
- Splunk Cloud
- Elasticsearch (ELK Stack)
- Kibana
- Prometheus
- Grafana
- Grafana Tempo
- OpenTelemetry
- Kafka
- Terraform
- Kubernetes
- Docker
- Linux
- Python, Go, Ruby, Bash
- AWS, Azure, GCP
- Ansible
- Consul
Work Authorization Requirements
- Must be a U.S. Citizen or U.S. National.
- Work must be performed within the United States.
- Candidates must be eligible to support FedRAMP High and IL-5 environments.
Experience
- 7+ years of relevant experience in SRE, DevOps, or Platform Engineering.
- Experience leading observability initiatives in enterprise-scale environments is highly preferred.
This is an excellent opportunity to lead critical observability initiatives and build scalable, resilient platforms supporting mission-critical cloud infrastructure across the United States.
Lead Site Reliability Engineer – Observability · Info Way Solutions LLC