Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
HI

Senior Observability Engineer

Han IT Staffing
Location not stated
Hybrid
Senior
4 weeks ago
  • Elastic Stack
  • Kubernetes
  • Microservices
  • Elasticsearch
  • Logstash
  • Kibana
  • Elastic
  • OpenShift
  • CI/CD
  • Python
  • Groovy
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Job Title: Senior Observability Engineer (ESS Platform SME)

Role Overview:
We are seeking ahighly experienced Senior Observability Engineer with deep expertise in ESS (Elastic Stack) to lead and accelerate the development of enterprise-grade observability capabilities across mission-critical applications.
This role requires a hands-on SME who candesign, build, and scale observability dashboards, APM, tracing, and monitoring solutions exclusively within ESS. The candidate will play a key role in transforming current monitoring into aproactive, intelligent, and scalable observability ecosystem.
This is ahigh-impact, fast-paced engagement (target< 6 months) requiring ownership, technical depth, and execution excellence.
Key Responsibilities:
ESS Observability Architecture & Implementation
  • Design and implementend-to-end observability solutions using ESS (Elastic Stack).
  • Build acentralized observability layer covering all MF applications.
  • Ensureblock-level aggregation with drill-down to:
  • Application-level metrics
  • APM traces
  • Logs and events
  • Service dependencies
Dashboard Engineering (Critical Priority)
  • Develop and scale alarge backlog of ESS dashboards, including but not limited to:
  • Cluster Health (OCP/K8s)
  • API & APM Dashboards
  • Service Health & Dependency Monitoring
  • Pod Status / Restart / Scaling Metrics
  • HTTP Status Analytics (200/400/500 trends)
  • Transaction Processing Metrics
  • Infra Metrics (CPU, Memory, Disk, Network)
  • Synthetic Monitoring & Availability
  • Buildintuitive, drill-down dashboards from MF Block → Service → Application level.
APM, Tracing & Monitoring Expansion
  • Expand ESS-based:
  • Application Performance Monitoring (APM)
  • Distributed tracing
  • Real User Monitoring (RUM)
  • Synthetic monitoring
  • Enableend-to-end traceability across microservices.
Proactive Observability & Alerting
  • Design and implementsmart alerting rules:
  • Move from reactive → proactive detection
  • Reduce noise, improve signal quality
  • DefineSLOs, SLIs, and error budgets
  • Enhance anomaly detection and trend analysis
Collaboration & Leadership
  • Work closely with:
  • EOT Observability Team
  • Internal CDLs
  • Application teams
  • Act asESS Observability SME
  • Provideguidance, standards, and best practices
Required Skills & Experience:
  • Strong hands-on experience with ESS (Elastic Stack):
  • Elasticsearch
  • Logstash
  • Kibana
  • Beats / Elastic Agent
  • Elastic APM
  • Proven experience buildingenterprise-scale observability dashboards in ESS
  • Deep understanding of:
  • Microservices architecture
  • Kubernetes / OpenShift (OCP)
  • Experience with:
  • APM, distributed tracing, logging, metrics correlation
  • Ability to designmulti-layer observability (infra → platform → app)
Strongly Preferred:
  • Experience with:
  • Synthetic monitoring tools integrated with ESS
  • Real User Monitoring (RUM)
  • Service maps and dependency graphs
  • Knowledge of:
  • CI/CD observability integration
  • Alerting frameworks within Elastic
  • Scripting: Python / Shell / Groovy (nice to have)
Soft Skills:
  • Strong ownership mindset
  • Ability to work under aggressive timelines
  • Excellent problem-solving skills
  • Clear communication with technical and non-technical teams
Success Criteria (First 3–6 Months):
  • Deliverenterprise-grade ESS observability dashboards
  • Achievefull MF application visibility
  • Implementend-to-end APM + tracing coverage
  • Establishproactive alerting framework
Additional Notes:
  • Candidatemust be an ESS expert — alternative tools experience alone will not be sufficient.
  • This is ahigh-priority, business-critical role with immediate impact expectations.

Senior Observability Engineer · Han IT Staffing

Auto apply with Likeremote