Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TC

Grafana Observability SME

TechDigital Corporation
πŸ‡ΊπŸ‡Έ United States
On-site
Senior
3 months ago
  • Grafana
  • Mimir
  • Loki
  • OpenTelemetry
  • Java
  • .NET
  • Python
  • Node.js
  • eBPF
  • Cilium
  • ServiceNow
  • AIOps
  • AWS
  • Azure
  • Linux
  • Windows
  • GitOps
  • Terraform
  • SolarWinds
  • Prometheus
  • JSON
  • Winston
  • FinOps
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Top Skills:
1. Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
2. Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
3. OpenTelemetry practitioner β€” OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
4. eBPF-based auto-instrumentation experience with Beyla (or equivalent β€” Pixie, Cilium Tetragon) in a production context.
5. Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
6. Multi-environment hosting fluency β€” on-prem, AWS, Azure β€” and Linux/Windows host agent deployment at scale.
7. Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
8. Excellent written communication β€” solution architecture documents, runbooks, and stakeholder-facing status reporting.

Role Summary
Own the end-to-end technical design, build, and operationalization of the Grafana Cloud observability platform for a 50-application estate spanning Java, .NET, Go, Python, and Node.js workloads hosted across on-premises data centres, AWS, and Azure. The SME serves as the senior technical authority across all eight in-scope Grafana Cloud modules and is accountable for instrumentation strategy, alerting design, dashboarding standards, and integration into ServiceNow ITOM via native Event Management. Scope is application-level observability only β€” server and network health remain on SolarWinds, and URL/synthetic monitoring remains on Uptrends.

Key Responsibilities
β€’ Platform architecture and configuration across all eight in-scope Grafana Cloud modules: Grafana 12 (visualization), Mimir (metrics, 13-month retention), Loki (logs), Tempo (distributed tracing via OTLP), Alloy (telemetry collection agent), Beyla (eBPF zero-code auto-instrumentation), Application Observability (OTel-native APM), and Unified Alerting.
β€’ Tenancy and access design β€” organizations, folders, teams, role-based access control, dashboard variables, template links, and annotations.
β€’ Application instrumentation strategy by technology stack: Beyla eBPF as the default zero-code path for Simple and Medium apps; OpenTelemetry SDKs/agents (Java, .NET, Go, Python, Node.js) for Complex apps requiring deeper traces and custom metrics; JMX Exporter, prometheus_client, and runtime-specific exporters where stack-appropriate.
β€’ Log pipeline engineering via Alloy β€” structured JSON, Log4j/Logback, Serilog, NLog, Windows Event Log, Winston, Pino, loguru β€” with parsing rules tuned per stack and LogQL-based dashboards and alerts.
β€’ Alerting design β€” PromQL/LogQL/TraceQL rules, severity taxonomy, grouping, routing, and notification policies. Build a low-noise, actionable alert feed; tune thresholds iteratively with application owners.
β€’ Single Pane of Glass β€” design and deliver a tiered SPoG that surfaces Grafana application telemetry alongside contextual links to SolarWinds and Uptrends.
β€’ Business Dashboards and Reporting β€” partner with the Dashboard Lead to define KPI taxonomy and ensure dashboard-as-code patterns and version control.
β€’ ServiceNow ITOM integration β€” co-own the design and review of Grafana β†’ ServiceNow Event Management (native inbound integration) flow: event allow-list governance ( "deny by default "), enrichment, deduplication, AIOps correlation, automated incident creation with severity mapping and assignment group rules, CMDB CI attachment, and ServiceNow-as-master incident state.
β€’ Quality assurance authority across all technical deliverables β€” solution architecture document, instrumentation runbooks, dashboard and alert library, integration test results.
β€’ Phased delivery execution β€” Mobilise & Client β†’ Application Foundation (ML1) β†’ Onboarding of 40 Simple apps (ML2) β†’ Medium/Complex apps + ITOM Integration (ML2β†’3) β†’ SPoG, Dashboards & Reporting (ML3β†’4) β†’ Stabilisation, KT, and post-deployment support (ML4).
β€’ Knowledge transfer β€” produce platform operating procedures and conduct structured handover to the client's run team.

Required Skills & Experience
β€’ 7+ years in observability/monitoring engineering with deep, recent hands-on Grafana Cloud experience (not just OSS Grafana).
β€’ Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
β€’ Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
β€’ OpenTelemetry practitioner β€” OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
β€’ eBPF-based auto-instrumentation experience with Beyla (or equivalent β€” Pixie, Cilium Tetragon) in a production context.
β€’ Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
β€’ Multi-environment hosting fluency β€” on-prem, AWS, Azure β€” and Linux/Windows host agent deployment at scale.
β€’ Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
β€’ Excellent written communication β€” solution architecture documents, runbooks, and stakeholder-facing status reporting.

Nice to Have
β€’ Grafana Certified Professional or equivalent vendor certification.
β€’ Prior experience in a regulated utility, energy, or critical-infrastructure environment.
β€’ Familiarity with SolarWinds and Uptrends (sufficient to design clean boundaries with retained tooling, not to administer them).
β€’ Experience with ServiceNow CSDM and Service Mapping governance.
β€’ Exposure to FinOps for observability β€” cardinality control, log volume management, retention tuning in Mimir/Loki.

Out of Scope for This Role
β€’ Server health and network monitoring (owned by SolarWinds).
β€’ URL/synthetic endpoint monitoring (owned by Uptrends).
β€’ ServiceNow ITSM workflow ownership β€” incident lifecycle remains with the client's ITSM/ITOM team; this role designs the integration, not the downstream process.

Grafana Observability SME Β· TechDigital Corporation

Auto apply with Likeremote