Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
PL

Platform (Dev) Engineer - Remote

PeopleNTech LLC
πŸ‡ΊπŸ‡Έ United States
Remote
Mid level
2 months ago
  • Grafana
  • Mimir
  • Loki
  • Prometheus
  • Kubernetes
  • JSON
  • YAML
  • SNMP
  • REST API
  • OpenTelemetry
  • Incident Management
  • ServiceNow
  • AIOps
  • Terraform
  • Ansible
  • Configuration Management
  • GitLab CI/CD
  • Python
  • GitOps
  • IaC
  • Helm
  • PostgreSQL
  • GitLab
  • CI/CD
  • Zabbix
  • Secrets Management
  • CyberArk
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Platform (Dev) Engineer - Remote
Indent ID: SF_OP_205310-2-1
Sell Rate: 70/hr
Remote / WFO / Hybrid:Remote. Ideally, Central/Mountain/Pacific TZ hours availability 8am onwards.
Hire Type (FTE/Contract): Contract
Project Duration: 12 Months
Expected Start Date: ASAP

Platform (Dev) Engineer

We are seeking a mid-level Platform Engineer (5+ years of relevant experience) to configure, instrument, and operationalize Flexential's enterprise Observability platform. This role is focused on the LGTM stack β€” installing, configuring, and maintaining Grafana, Mimir, Loki, Tempo, Prometheus, and OTel collectors (or Grafana Alloy) β€” and on building the monitoring and observability workflows that support 40+ data center facilities and a broad range of infrastructure and application teams.

Responsibilities
β€’ Install, configure, and administer the full LGTM stack (Loki, Grafana, Tempo, Mimir) in Kubernetes environments
β€’ Build and modify Grafana dashboards as code (JSON), including pre-built dashboards for infrastructure, application, and facility teams
β€’ Author and maintain Prometheus scrape configurations (YAML) for diverse device and application targets including SNMP, REST API, and OpenTelemetry sources
β€’ Help define and implement end-to-end Major Incident Management (MIM) workflows integrating Alertmanager, ServiceNow, and CMDB
β€’ Help identify and implement AIOps capabilities including anomaly detection, event correlation, and alert noise reduction
β€’ Build and maintain Terraform and Ansible configurations supporting platform deployments and configuration management
β€’ Develop GitLab CI/CD pipelines for automated delivery of platform configuration changes to Central and Edge stacks
β€’ Help onboard infrastructure and application teams to the platform, providing YAML templates, scrape profiles, and dashboard standards
β€’ Contribute to observability platform runbooks, onboarding guides, and architecture documentation

Required Skills
Python and YAML β€” 3+ years of automation scripting, configuration authoring, and pipeline development
β€’ Terraform and Ansible β€” 2+ years in a GitOps model; Infrastructure-as-Code (IaC) and configuration management experience required
β€’ Kubernetes and Helm β€” 2+ years deploying and managing production workloads; LGTM stack experience in Kubernetes strongly preferred
β€’ PostgreSQL (or similar)β€” 1+ years of administration and platform integration experience
β€’ GitLab (or similar) β€” 2+ years of CI/CD pipeline development and GitOps workflow management
LGTM stack (Loki, Grafana, Tempo, Mimir) β€” Strong experience with hands-on installation, configuration, and administration required; this is a core requirement
β€’ Grafana β€” Experience authoring dashboard JSON, configuring data sources, and building alerting rules
β€’ Prometheus β€” 2+ years writing scrape configs, configuring service discovery, and managing remote_write pipelines
β€’ OpenTelemetry β€” 1+ years configuring collectors, working with OTLP pipelines, and understanding SDK instrumentation concepts
β€’ MIM workflow design β€” Experience designing and implementing end-to-end incident management workflows based on ITSM standards


Preferred Skills
β€’ Prior experience integrating observability platforms with ServiceNow for alert-toticket automation
β€’ Familiarity with Zabbix for device auto-discovery and inventory management at scale
β€’ Experience with AIOps tooling for anomaly detection and event correlation
β€’ Experience with secrets management integrations using CyberArk and Conjur

Platform (Dev) Engineer - Remote Β· PeopleNTech LLC

Auto apply with Likeremote