Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
IW

Lead Site Reliability Engineer – Observability (US Remote)

Info Way Solutions LLC
Location not stated
Remote
Staff / Principal
2 months ago
  • Splunk
  • Elasticsearch
  • Grafana
  • OpenTelemetry
  • Prometheus
  • Kibana
  • Terraform
  • Configuration Management
  • Devops
  • IaC
  • Python
  • Ruby
  • Bash
  • Kubernetes
  • AWS
  • Azure
  • GCP
  • Ansible
  • CI/CD
  • FedRAMP
  • Docker
  • Linux
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Lead Site Reliability Engineer – Observability (US Remote)
Level: L2
Location: Remote, prefer PST hours
Visa Requirements: UC Citizenship
About the Role
Join our Observability team responsible for designing, building, and operating enterprise
platforms for logging, metrics, tracing, and alerting across large-scale cloud infrastructure.
You'll lead initiatives that improve reliability, scalability, and operational excellence.
Key Responsibilities
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including
Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and
OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and
trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring
solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL,
Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Required Qualifications
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo,
distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern
observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
Preferred Qualifications
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines,
and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana,
Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes,
Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
US Requirement
The successful applicant may perform work in FedRAMP High or IL-5 environments and
therefore must be a U.S. Person (U.S. citizen or U.S. national). Work must be performed
from within the United States.
Keywords
SRE, Site Reliability Engineering, Observability, Splunk, Splunk Enterprise, Splunk Cloud,
Elasticsearch, ELK, Kibana, Prometheus, Grafana, Tempo, Grafana Tempo, Distributed
Tracing, OpenTelemetry, Traces, Kafka, Terraform, Kubernetes, Linux, DevOps.

Lead Site Reliability Engineer – Observability (US Remote) · Info Way Solutions LLC

Auto apply with Likeremote