Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
L

Specialist - Cloud & Infra Management

LTM
Location not stated
1 month ago
  • Windows
  • Linux
  • Unix
  • AWS
  • Azure
  • GCP
  • Incident Management
  • Change Management
  • ITIL
  • AIOps
  • AI/ML
  • Dynatrace
  • Datadog
  • SolarWinds
  • Zabbix
  • Nagios
  • Prometheus
  • Grafana
  • Elastic
  • New Relic
  • OpenTelemetry
  • ServiceNow
  • Jira Service Management
  • Configuration Management
  • PRINCE2
  • AWS Cloud
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Role Summary

TheTools Lead - Observability, Monitoring & Service Management will lead the strategy, architecture, implementation, optimization, and governance of enterprise observability and monitoring platforms across infrastructure, cloud, network, application, and digital services environments.

The role is responsible for driving tool standardization, platform modernization, vendor management, operational excellence, and service integration initiatives to enhance visibility, reliability, performance, and business outcomes.

This position will own enterprise monitoring capabilities across on-premises, cloud, and hybrid environments while collaborating with delivery, operations, engineering, service management, and customer-facing teams.

Key Responsibilities Observability & Monitoring Leadership
  • Define and drive the enterprise observability and monitoring strategy.
  • Lead the implementation, optimization, and governance of monitoring platforms across infrastructure and cloud environments.
  • Establish monitoring standards, KPIs, SLAs, dashboards, ingestion frameworks, and operational best practices.
  • Drive monitoring maturity initiatives, including:
    • Event correlation
    • Noise reduction
    • Predictive analytics
    • Operational automation
Infrastructure Monitoring

Manage monitoring solutions covering:

  • Windows, Linux, and Unix servers
  • Network devices
  • Storage systems
  • Databases
  • Virtualization platforms
  • Cloud services (AWS, Azure, GCP)
  • Middleware platforms

Key responsibilities include:

  • Ensuring end-to-end visibility across enterprise infrastructure environments.
  • Driving proactive monitoring and performance management.
  • Supporting capacity planning and availability management initiatives.
Application Performance Monitoring (APM)
  • Collaborate with application teams to implement Application Performance Monitoring (APM) solutions.
  • Enable monitoring of:
    • Application availability
    • Performance metrics
    • Transaction tracing
    • End-user experience
    • Service health
  • Support root cause analysis and application performance optimization initiatives.
Tool Governance & Vendor Management
  • Evaluate, onboard, and govern strategic monitoring and observability vendors.
  • Manage tool licensing, renewals, budgets, contracts, and optimization opportunities.
  • Drive enterprise-wide tool rationalization and standardization across business units.
  • Ensure vendor performance aligns with organizational objectives and service commitments.
ITSM & Service Integration
  • Integrate monitoring tools with ITSM platforms to enable automated:
    • Incident management
    • Event management
    • Change management workflows
  • Improve operational efficiency through automation and service orchestration.
  • Support ITIL-based service management processes and continuous service improvement initiatives.
Automation & Innovation
  • Drive adoption of AIOps, observability automation, self-healing capabilities, and event correlation frameworks.
  • Identify opportunities to leverage AI/ML for proactive monitoring and operational intelligence.
  • Promote modern observability practices across engineering and operations teams.
  • Champion innovation through automation-first approaches.
Stakeholder Management
  • Partner with delivery teams, operations teams, architects, engineering teams, customers, and senior leadership.
  • Provide:
    • Executive dashboards
    • Governance reports
    • Risk assessments
    • Strategic recommendations
  • Act as the Subject Matter Expert (SME) and escalation point for monitoring platform strategy and operations.
  • Mentor and guide junior team members to build organizational capability.
Required Experience
  • 5-7 years of overall IT experience.
  • Minimum 5 years of experience in Monitoring, Observability, and Enterprise Tools leadership roles.
  • Strong experience managing enterprise monitoring platforms and observability ecosystems.
  • Proven leadership in large-scale implementation, transformation, and modernization programs.
  • Experience in vendor management, governance, budgeting, and stakeholder engagement.
  • Strong understanding of:
    • IT Infrastructure
    • Cloud Operations
    • Site Reliability Engineering (SRE)
    • IT Operations
    • IT Service Management (ITSM)
Primary Skills (Mandatory) Observability & Infrastructure Monitoring

Hands-on experience with one or more of the following platforms:

  • OpsRamp
  • LogicMonitor
  • ScienceLogic
  • Dynatrace
  • Datadog
  • SolarWinds
  • Zabbix
  • Nagios
  • Microsoft SCOM
  • PRTG
  • Prometheus
  • Grafana
  • Elastic Observability
Infrastructure Technologies
  • Windows Monitoring
  • Linux Monitoring
  • Network Monitoring
  • Storage Monitoring
  • Database Monitoring
  • Virtualization Monitoring
  • Public Cloud Monitoring (AWS, Azure, GCP)
Leadership & Governance
  • Program Management
  • Vendor Management
  • Governance & Compliance
  • Stakeholder Management
  • Service Delivery Leadership
  • Financial Management
  • Contract Management
Secondary Skills (Preferred) Application Monitoring

Experience with:

  • Dynatrace
  • AppDynamics
  • New Relic
  • Datadog APM
  • Elastic APM
  • OpenTelemetry
ITSM & Service Management

Experience with:

  • ServiceNow
  • BMC Remedy
  • Jira Service Management
  • Ivanti
  • ManageEngine
Automation & AIOps
  • AIOps Platforms
  • Event Correlation
  • Runbook Automation
  • Workflow Automation
  • Configuration Management Tools
Preferred Certifications

Monitoring & Observability

  • OpsRamp Professional Certification
  • ScienceLogic Professional Certification

Project & Service Management

  • PRINCE2
  • ITIL® 4 Foundation / Managing Professional

Good-to-Have Certifications

  • Dynatrace Associate / Professional
  • Datadog Certification
  • ServiceNow Certification
  • AWS Cloud Certification
  • Microsoft Azure Certification
  • SRE Foundation Certification

Specialist - Cloud & Infra Management · LTM

Auto apply with Likeremote