Senior Site Reliability Engineer - OpenTelemetry
- OpenTelemetry
- OpenSearch
- Grafana
- Incident Response
We're looking for aSenior Site Reliability Engineer โ Observability to join our team in London, United Kingdom in a hybrid working mode. You will be part of the Production Engineering โ Observability team, driving the strategic initiative to implement and expand a modern observability platform built on OpenTelemetry and OpenSearch. This program focuses on enhancing monitoring, resilience and operational stability for critical FIC trading systems by enabling faster incident detection, reducing outage duration and providing actionable insights across the technology estate. This is a hands-on engineering role where youโll define observability standards and deliver enterprise-scale solutions while promoting best practices for operational excellence.
Responsibilities
- Gather requirements and perform analysis of existing monitoring and observability platforms
- Define observability standards, telemetry strategies and alerting frameworks
- Implement OpenTelemetry-based instrumentation and OpenSearch solutions across applications and infrastructure
- Design dashboards, analytics and reporting to improve transparency and operational efficiency
- Develop automation tools and processes to enhance observability and reduce manual overhead
- Integrate observability frameworks with enterprise monitoring platforms such as Geneos
- Provide documentation and operational handover to ensure long-term sustainability
- Apply Site Reliability Engineering principles to drive stability, scalability and incident reduction
Requirements
- Proven experience as Senior SRE or Observability Engineer implementing enterprise-scale observability solutions
- Strong expertise in OpenTelemetry including instrumentation and telemetry pipelines
- In-depth knowledge of OpenSearch for architecture, data indexing, optimisation and analytics
- Experience developing dashboards and alerts using Grafana
- Familiarity with enterprise monitoring tools such as Geneos and related observability technologies
- Practical understanding of Site Reliability Engineering practices and automation approaches
- Background in high-availability or mission-critical environments; financial services experience is highly desirable
Nice to have
- Knowledge of anomaly detection, alert correlation, and incident response automation
- Exposure to observability in cloud-native or hybrid architectures
Senior Site Reliability Engineer - OpenTelemetry ยท EPAM Systems