
Operations Specialist
- Datadog
- triage
- Incident Response
- Devops
- AWS
- Azure
- Kubernetes
- Dynatrace
- Grafana
- Prometheus
- Apigee
- AI
- Unix
- Linux
- GCP
- Docker
- Python
- Bash
- PowerShell
- Jira
- ServiceNow
- Active Directory
- FTP
- DNS
- SSH
- CI/CD
- Jenkins
- GitHub Actions
- Azure DevOps
- ITIL
- Change Management
- Atlassian
- Confluence
About the Role:
OEC is in the middle of a multi-year journey to modernise its infrastructure and applications, transforming how we build, deploy, and operate software at global scale. The Operations Specialist (L2) is a core member of the Monitoring Team within the Enterprise Operations Team, responsible for the configuration, maintenance, and continuous improvement of OEC’s observability stack, primarily powered by Datadog.
Â
In this role you will own L2 alert triage, incident response, dashboard development, and service onboarding — contributing to the team’s tooling standards and best practices. You will also handle day-to-day server monitoring, infrastructure support, and service management tasks that keep OEC’s global platforms running reliably.
Key Responsibilities & Duties (essential to the job)
Monitoring & alerting
- Design, configure, and maintain Datadog monitors, composite alerts, and notification channels for production and non-production environments.
- Own the L2 alert triage process — investigate, diagnose, and resolve escalated alerts from L1, ensuring timely root cause identification.
- Tune alert thresholds and suppression rules to minimise noise and reduce false positive rates.
- Manage SLO (Service Level Objective) definitions, tracking, and reporting across assigned services.
- Participate in the on-call rotation for critical monitoring alerts and P1/P2 incident response.
Dashboards & observability
- Build and maintain Datadog dashboards covering infrastructure health, application performance, log analytics, and business KPIs.
- Configure and manage Datadog APM (Application Performance Monitoring), distributed tracing, and error tracking for key services.
- Implement and manage log ingestion pipelines, parsing rules, and log-based monitors.
- Develop and maintain Synthetic monitors (API and browser tests) for uptime and user experience validation.
Â
Service onboarding & integrations
- Onboard new services and infrastructure components into the Datadog monitoring framework in collaboration with DevOps, Cloud, and Engineering teams.
- Configure and maintain Datadog integrations with cloud platforms (AWS, Azure, Rackspace), Kubernetes, containerised workloads, and third-party tools.
- Support the migration of legacy monitoring tooling (SCOM, Dynatrace, Grafana, Prometheus, Pingdom, Apigee, Redgate, IDERA) into Datadog.
Incident response & runbooks
- Act as L2 escalation point during incidents — perform deep-dive investigations using metrics, logs, traces, and dashboards.
- Create, maintain, and improve runbooks for common alert scenarios, incident response procedures, and post-incident remediation steps.
- Contribute to post-incident reviews, documenting root cause findings and identifying monitoring gaps to prevent recurrence.
- Alerts system owners, stakeholders, and management of degraded system status and Priority 1 and Priority 2 incidents; issues tickets for incident and problem escalation.
AI & automation
- Leverage Datadog AI capabilities including Watchdog, anomaly detection, outlier detection, and Bits AI to improve proactive issue detection.
- Implement forecast-based monitors for capacity management (CPU, memory, disk).
- Use AI-based alert correlation and deduplication features to reduce alert fatigue.
- Contribute to automation of repetitive monitoring tasks using Datadog’s API and scripting tools.
Standards & governance
- Adhere to and activelycontribute to monitoring standards, tagging strategies, and naming conventions.
- Participate in regular alert quality reviews, dashboard audits, and SLO compliance checks.
- Maintain accurate documentation of monitoring configurations, integrations, and architectural decisions.
- Creates and maintains knowledge articles to be used internally and/or externally for training, best practices, solutions, or processes relating to applications, environments, and related technologies.
Infrastructure operations & support
- Analyses and troubleshoots Microsoft and UNIX/Linux server configurations and processes.
- Performs moderately complex database administration.
- Monitors data centre networks, infrastructure bandwidth, application health, servers, and other infrastructure; coordinates and communicates with vendors.
- Diagnoses and researches (using knowledge base) application incidents, monitoring alerts, and service requests; provides assistance and guidance to associate operations specialists.
- Adheres to all incident and service request processes and procedures in accordance with established Service Level Agreements (SLAs).
- Diagnoses, researches, and resolves Level-2 technical hardware and software incidents, monitoring alerts, and service requests.
- Works on ad-hoc projects to support Infrastructure Engineering or other departments, as requested.
- Demonstrates a flexible and adaptable approach to work and adjusts to shifts in priorities as the needs of the business change.
- Collaborates with DevOps, Cloud, and application teams to align monitoring coverage with service requirements.
Education
A bachelor’s degree from an accredited college or university in Computer Science, Information Technology, Engineering, or a related discipline is required. In the absence of a degree, equivalent work experience directly related to the key responsibilities of the role will be considered as a substitute for the degree.
Experience, Skills and Key Competencies
At least 3–5 years of experience in an infrastructure, DevOps, SRE, or monitoring engineering role is required, with hands-on experience in Datadog and working knowledge of cloud infrastructure and containerised environments.
Â
Must also be able to demonstrate the following skills and abilities:
Â
Â
Technical skills
- Hands-on experience with Datadog — monitors, dashboards, APM, log management, SLOs, and synthetic monitoring.
- Solid understanding of cloud infrastructure (AWS required; Azure or GCP a plus) and containerised environments (Docker, Kubernetes).
- Familiarity with observability concepts: metrics, traces, and logs.
- Experience with at least one scripting language (Python, Bash, or PowerShell) for automation and tooling.
- Working knowledge of monitoring and alerting tools such as Prometheus, Grafana, Dynatrace, or similar platforms.
- Understanding of ITSM concepts and ticketing workflows (Jira, JSM, or ServiceNow).
- Good working knowledge of Microsoft and UNIX/Linux server configurations, Active Directory, and internet protocols (HTTP, FTP, DNS, IP, SSH).
Soft skills & competencies
- Strong analytical and problem-solving skills with the ability to diagnose complex infrastructure issues under pressure.
- Clear written and verbal communication skills; able to produce runbooks and incident summaries for both technical and non-technical audiences.
- Collaborative team player who can work effectively across Engineering, DevOps, and Operations disciplines.
- Self-motivated with a continuous improvement mindset and a proactive approach to identifying and resolving issues before they escalate.
- Flexible and adaptable approach to work; able to adjust to shifts in priorities as the needs of the business change.
Preferred Qualifications
- Datadog Fundamentals or Datadog APM certification (or actively working towards one).
- At least 1 year of experience with CI/CD pipelines (Jenkins, GitHub Actions, or Azure DevOps).
- Knowledge of network monitoring and distributed systems architecture.
- Exposure to ITIL or ITSM frameworks (Incident, Problem, and Change Management).
- Familiarity with Atlassian tools (Jira, JSM, Confluence) for documentation and ticket management.
Special Position Requirements
- Must be able to read, write, understand, and speak fluent English.
- Must be able to work flexible shifts including holidays and weekends, to provide support across locations and time zones.
- Must be able to participate in an on-call rotation for critical monitoring alerts and P1/P2 incident response.
Â
Â
Â
Â
Â
Â
Operations Specialist · OEConnection