Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Senior Platform Engineer (Observability)

OneMain General Services Corporation
🇺🇸 United States
On-site
Senior
1 day ago
  • Grafana
  • Elastic Stack
  • AWS
  • CloudWatch
  • Azure Monitor
  • OpenTelemetry
  • InfluxDB
  • Bash
  • PowerShell
  • Python
  • JavaScript
  • Agile
  • Configuration Management
  • IaC
  • Ansible
  • Terraform
  • REST API
  • JSON
  • ServiceNow
  • Azure
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We’re seeking aSenior Monitoring Engineer to join a high‑performing Monitoring Engineering team in a fast‑paced finance technology organization. You’ll design, develop, and maintain monitoring and observability solutions that keep core applications and infrastructure healthy and visible. In close partnership with application, platform, and development teams, you will implement alerting systems, dashboards, correlations, and automation—driving reliability, reducing MTTR, and elevating operational awareness.

Critical thinking, system analysis, and proactive troubleshooting are essential to success in this role.

Key Responsibilities 

Design, Build, and Maintain Monitoring & Observability Solutions

  • Develop and maintaininstrumentation, telemetry, and alerting for the Enterprise Monitoring Center using industry‑leading tools, such as:
    • Grafana
    • OpsRamp
    • AppDynamics
    • Elastic Stack
    • BigPanda
    • AWS CloudWatch
    • Azure Monitor
  • ImplementObservability best practices, ensuring comprehensive coverage ofmetrics, logs, and traces across critical systems.
  • Integrate and manageOpenTelemetry for distributed tracing and telemetry data collection, enabling end‑to‑end visibility of business‑critical transactions.

Collaboration & Project Participation

  • Collaborate with application development teams todefine and document observability requirements for each project or release.
  • Participate in complex initiatives, ensuring accurate and actionable monitoring and tracing are in place for every step of business‑critical workflows.

Alerting & Escalation Process

  • Define and maintainstandardized alert payloads per engineering guidelines, ensuring alerts areactionable.
  • Partner with Level 2 and Level 3 support teams to reflect process changes in monitoring dashboards.
  • Maintain and optimizethresholds, ensuring seamlessescalations viaBigPanda as the central alert hub.

Dashboard Creation & Maintenance

  • Create and maintainintuitive, actionable dashboards for the Enterprise Monitoring Center and other finance teams.
  • Ensure dashboards are effectivelymonitored by Level 1 teams, presenting clear, actionable data thatreduces MTTR.

System Validation, Documentation & Automation

  • Develop and maintainautomation scripts to enhance monitoring efficiency and improve team quality of life.
  • Proactively identifyprocess improvements and learning opportunities; drivecontinuous improvement.

Automation & Quality‑of‑Life Improvements

  • Contribute to theautomation of monitoring, alerting, and operational tasks to streamline workflows and improve overall system reliability.

Qualifications

Education

Bachelor’s in Computer Science, IT, or related field.

Experience

  • Minimum 4 years in a technology organization, with≥1 year hands‑onengineering experience inmonitoring or production operations.

Required Skills

  • Strong experiencedeveloping instrumentation and alerting forlarge, complex environments.
  • Expertise in≥4 of the following:OpsRamp, Grafana, AppDynamics, Elastic Stack, InfluxDB, BigPanda, and other monitoring solutions.
  • Hands-on experience with Observability concepts and frameworks, includingmetrics, logs, and traces.
  • Working knowledge of OpenTelemetry for distributed tracing and telemetry data collection.
  • Experience withdashboard creation,alert management, andtool configuration.
  • Excellentverbal and written communication—able to present complex technical issues to both technical and non‑technical stakeholders.
  • Strongproblem‑solving and troubleshooting inhigh‑pressure environments.
  • Ability toprioritize and manage multiple tasks in adeadline‑driven setting.
  • Proven collaboration withcross‑functional teams inlarge, complex IT environments.
  • Experience withscripting (e.g.,Bash,PowerShell) and proficiency inone programming language (e.g.,Python,C family,JavaScript).
  • Experience designing and implementingscalable, reliable monitoring solutions.
  • Experience with agile software development methodologies
  • Familiar with problem diagnosis; performance tuning; capacity planning and configuration management across the stack via continuous improvement.

Preferred Qualifications

  • Experiencequerying, manipulating, and visualizing time‑series data.
  • Familiarity withInfrastructure as Code tools (e.g.,Ansible,Terraform).
  • Strong understanding of how to createactionable, digestible visualizations forLevel 1 monitoring teams.
  • Working knowledge ofREST APIs,JSON, andServiceNow.
  • Experience withcloud monitoring—particularlyAWS orAzure.

OneMain Holdings, Inc. is an Equal Employment Opportunity (EEO) employer. Qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship status, color, creed, culture, disability, ethnicity, gender, gender identity or expression, genetic information or history, marital status, military status, national origin, nationality, pregnancy, race, religion, sex, sexual orientation, socioeconomic status, transgender or on any other basis protected by law.

Senior Platform Engineer (Observability) · OneMain General Services Corporation

Auto apply with Likeremote