
Principal Observability & Reliability Architect
AHEAD
🇺🇸 United States
Hybrid
Staff / Principal
11 months ago
$180,000 – $240,000 / year
- Dynatrace
- OpenTelemetry
- AI
- Incident Response
- AIOps
- Kubernetes
- Prometheus
- Grafana
- AWS
- Azure
- GCP
- Terraform
- Ansible
- Python
- ITIL
- Devops
- New Relic
- Datadog
- Splunk
- Fluent Bit
- CI/CD
- IaC
- ServiceNow
- Jira Service Management
- PagerDuty
- Opsgenie
- FinOps
- Pension
11 months ago
AHEAD is seeking a Principal Observability & Reliability Architect to join our Observability practice within Intelligent Operations. You are the senior technical authority and client advisor for observability and reliability: an expert in Dynatrace at enterprise scale, fluent in open standards and adjacent platforms, and accountable for the outcomes your programs deliver in reliability, efficiency, cost, and adoption. You will architect enterprise observability solutions, lead the SRE practices that turn telemetry into dependable services, advise executive stakeholders, serve as escalation point for delivery teams, and grow the practice through pre-sales, offerings, reusable content, and mentorship. This is a full-time, remote role based in the United States, with occasional travel (0 to 15%) based on engagement needs.
Responsibilities
- Architect Dynatrace at enterprise scale: multi-tenant and hybrid designs, ActiveGate topology, OneAgent and OpenTelemetry instrumentation strategy, consumption and licensing governance, tagging and ownership models, and AI-workload observability.
- Design end-to-end observability architectures across monitoring, logging, metrics, tracing, telemetry pipelines, alerting, event correlation, and service visibility in hybrid and multi-cloud environments.
- Establish and mature SRE practices with clients: SLIs and SLOs, error budgets, production readiness, incident response and postmortems, and reliability roadmaps tied to business impact.
- Lead assessment and advisory workshops that define use cases, maturity roadmaps, operating models, and adoption strategies, including AIOps and automation with Davis AI, alerting profiles, and workflows.
- Define standards for telemetry onboarding, naming, tagging, service ownership, access, dashboards, alert governance, runbooks, and operational handoff, and advise on telemetry governance: data quality, retention, sampling, cardinality, and cost.
- Lead modernization initiatives: tool and alert rationalization, telemetry strategy, migration to Dynatrace from legacy APM and monitoring platforms, and integration with ITSM, CMDB, event management, and automation platforms.
- Lead complex programs, own solution design and architectural review, articulate trade-offs, and act as escalation point for delivery teams.
- Advise client executives on platform strategy and value realization, and report on program health and outcomes.
- Provide architecture and quality oversight across engagements, intervening early where outcomes are at risk.
- Support pursuits as the technical expert: scoping, positioning, demonstrations, estimate validation, and client-facing technical narratives.
- Build reusable assets such as Dynatrace reference architectures, governance models, accelerators, and points of view, and contribute thought leadership through content, partner material, and conference speaking.
- Mentor architects and consultants across the practice, and maintain Dynatrace Professional certification plus professional-level certification on at least one additional platform.
Qualifications
- 8 or more years of hands-on experience in observability, APM, SRE, or related disciplines, including architecting enterprise-scale solutions across distributed systems and multi-cloud estates.
- 4 or more years of hands-on enterprise Dynatrace experience, including architecture and governance, OneAgent and Kubernetes deployment, Smartscape and PurePath, Grail and DQL, Davis AI, SLOs, workflows and automation, and ITSM integration; Dynatrace Professional certification held or attainable within six months.
- Applied SRE experience defining SLIs and SLOs, operating error budgets, running production readiness and incident reviews, and leading reliability programs that measurably reduce incidents and time to resolve.
- Working expertise in OpenTelemetry, Prometheus, and the Grafana ecosystem, and in public cloud monitoring services on AWS, Azure, or GCP.
- Strong knowledge of telemetry governance (routing, transformation, enrichment, retention, access, cost) and experience defining enterprise standards for dashboards, alerts, tagging, and service ownership.
- Expert knowledge of platform architecture, API integration patterns, and automation frameworks (Terraform, Ansible, Python, or similar).
- Strong consultative and executive-facing presence, with experience leading workshops and translating business needs into architecture and delivery plans.
- Demonstrated leadership mentoring technical teams; familiarity with ITIL, ITSM, and DevOps principles and with scoping, estimating, and change control in consulting delivery.
- Strong communication skills, attention to detail, and a self-starting work style; able to travel occasionally (0% to 15%).
Preferred Qualifications
- Both Dynatrace Professional certifications, or Dynatrace partner program experience.
- Migration experience from AppDynamics, New Relic, Datadog, Splunk, or legacy monitoring platforms to Dynatrace; LogicMonitor or other infrastructure monitoring experience is a plus.
- Telemetry pipeline tools such as OpenTelemetry Collector, Grafana Alloy, Fluent Bit, Kafka, Cribl, or Vector, plus Kubernetes, CI/CD, and infrastructure as code.
- Integration with ServiceNow, Jira Service Management, PagerDuty, Opsgenie, BigPanda, or xMatters.
- Published thought leadership, conference speaking, or ownership of a named offering or accelerator; relevant cloud, SRE, ITIL, or FinOps certifications are a plus.
Principal Observability & Reliability Architect · AHEAD