Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
DL

Observability

Diverse Lynx India
๐Ÿ‡ฎ๐Ÿ‡ณ India
On-site
4 months ago
  • triage
  • Incident Response
  • CI/CD
  • Kubernetes
  • Node.js
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Description:
Maintain and enhance observability systems, including alerting, dashboards, logs, traces, and metrics across the platform
Define, implement, and continuously monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to ensure platform reliability and performance
Support incident detection, triage, troubleshooting, and root cause analysis to minimize downtime and improve operational resilience
Design and maintain proactive alerting strategies to reduce noise and enable faster incident response
Integrate observability practices and tooling into the CI/CD pipelines to enable early detection of issues during build and deployment stages
8-10 years
Implement and manage observability solutions within a Kubernetes-based infrastructure, ensuring visibility into cluster, node, and application performance
Collaborate with development and platform teams to embed monitoring and reliability best practices across services
8-10 years
Analyze trends and metrics to identify performance bottlenecks and drive continuous improvement
Document observability standards, dashboards, alerts, and operational runbooks
Participate in post-incident reviews and contribute to reliability and availability improvements
Experience with large scale AD modernization or cloud identity transformation

Observability ยท Diverse Lynx India

Auto apply with Likeremote