Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
E

Observability Engineer

Evolutyz
Location not stated
Senior
2 months ago
  • Terraform
  • Python
  • FastAPI
  • OpenTelemetry
  • Datadog
  • IaC
  • CI/CD
  • REST API
  • Incident Response
  • Microservices
  • AWS
  • Azure
  • GCP
  • Kubernetes
  • Prometheus
  • Grafana
  • Splunk
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Summary:

We are seeking a highly skilled Senior Observability Engineer to design, implement, and optimize modern observability solutions across cloud-native environments. The ideal candidate will have strong expertise in infrastructure automation, backend development, and monitoring platforms, with hands-on experience in Terraform, Python, FastAPI, OpenTelemetry, and Datadog. This role will focus on building scalable observability frameworks, improving system reliability, and enabling proactive monitoring for critical applications and infrastructure.

Responsibilities:

  • Infrastructure as Code (IaC):
    • Design, build, and manage scalable cloud infrastructure using Terraform.
    • Develop reusable infrastructure modules and maintain infrastructure provisioning standards.
    • Implement and maintain CI/CD pipelines to automate infrastructure and application deployments.
    • Improve deployment reliability through automation and infrastructure best practices.
  • Backend Development:
    • Design, develop, and maintain RESTful APIs using Python and FastAPI.
    • Build backend services that support observability workflows, telemetry processing, and integrations.
    • Optimize services for performance, scalability, reliability, and maintainability.
    • Troubleshoot application and API performance issues.
  • Observability & Monitoring:
    • Design and implement observability solutions using OpenTelemetry for distributed tracing, metrics, and logging.
    • Configure and maintain Datadog dashboards, monitors, alerts, and Service Level Objectives (SLOs).
    • Develop monitoring strategies to improve system visibility and reduce incident response times.
    • Analyze telemetry data to identify performance bottlenecks and reliability issues.
    • Establish alerting thresholds and monitoring standards to support production operations.

Requirements:

  • Degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • 10 years of experience in cloud infrastructure, software engineering, or observability engineering.
  • Strong hands-on experience with Terraform and Infrastructure as Code practices.
  • Proficiency in Python and backend API development using FastAPI.
  • Experience implementing observability frameworks using OpenTelemetry.
  • Strong experience with Datadog, including dashboards, monitors, alerting, and SLO management.
  • Solid understanding of distributed systems, cloud-native architectures, and microservices.
  • Experience with CI/CD tools and deployment automation.
  • Strong troubleshooting, analytical, and problem-solving skills.

Preferred Skills:

  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud.
  • Knowledge of container orchestration platforms such as Kubernetes.
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, or Splunk.
  • Experience supporting high-availability production environments.

Benefits:

Details about benefits are not provided in the draft. Please specify benefits if applicable.

Observability Engineer · Evolutyz

Auto apply with Likeremote