Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
IW

SME SRE Observability

Info Way Solutions LLC
πŸ‡ΊπŸ‡Έ United States
On-site
Mid level
6 months ago
  • Devops
  • Prometheus
  • Grafana
  • Elastic Stack
  • Datadog
  • Splunk
  • AWS
  • Azure
  • GCP
  • Python
  • Bash
  • CI/CD
  • Docker
  • Kubernetes
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Job Title: SME – SRE Observability Engineer
Location: Minnesota (Onsite – 4 to 5 days/week)

Job Summary:
We are seeking an experiencedSubject Matter Expert (SME) in Site Reliability Engineering (SRE) with a strong focus on Observability. The ideal candidate will be responsible for designing, implementing, and optimizing observability frameworks to ensure high system reliability, performance, and scalability in a production environment.

Key Responsibilities:
  • Lead the design and implementation ofobservability solutions including metrics, logging, and tracing.
  • Act as an SME forSRE best practices, ensuring system reliability, availability, and performance.
  • Develop and maintain dashboards, alerts, and monitoring strategies.
  • Collaborate with development, DevOps, and infrastructure teams to improve system visibility.
  • Performroot cause analysis (RCA) and drive incident resolution.
  • Optimize system performance and reliability through proactive monitoring.
  • Implement automation to improve operational efficiency and reduce manual intervention.
  • Define and trackSLIs, SLOs, and SLAs.

Required Skills & Qualifications:
  • Strong experience inSite Reliability Engineering (SRE) concepts and practices.
  • Deep expertise inObservability tools (e.g., Prometheus, Grafana, ELK Stack, Datadog, Splunk, or similar).
  • Experience withcloud platforms (AWS, Azure, or GCP).
  • Proficiency inscripting/programming (Python, Bash, or similar).
  • Hands-on experience withmonitoring, alerting, and logging frameworks.
  • Strong troubleshooting and performance tuning skills.
  • Experience withCI/CD pipelines and automation tools.

Preferred Qualifications:
  • Experience working inhigh-availability, distributed systems.
  • Knowledge ofcontainerization and orchestration tools (Docker, Kubernetes).
  • Prior experience as anSRE SME or Lead.


SME SRE Observability Β· Info Way Solutions LLC

Auto apply with Likeremote