Senior Site Reliability Engineer
- AI
- LangGraph
- AWS
- CI/CD
- New Relic
- Terraform
- Opsgenie
- ServiceNow
- OpenTelemetry
- Machine Learning
- CloudWatch
- AWS Bedrock
- Kinesis
We are looking for aSenior Site Reliability Engineerto join our team in building an Enterprise Agent Development Platform β a production-grade, cloud-native ecosystem that enables engineering teams to define, orchestrate, deploy, and observe AI agents at scale. The platform standardizes agent development across the organization using LangGraph and Strands Agents on AWS AgentCore Runtime. This initiative spans agent framework design, runtime architecture, marketplace integration, CI/CD automation, and enterprise-grade observability β reducing agent development from months to days while enforcing consistent security, quality, and governance standards.
As a Senior Observability Engineer focused on New Relic Consolidation, you will own the multi-tenant New Relic consolidation layer, aggregating telemetry across the agent estate.
Responsibilities
- Own the multi-tenant New Relic consolidation layer aggregating telemetry across the entire agent estate
- Administer New Relic at scale, including management account setup and cross-account configuration
- Design and maintain NRQL cross-account queries aggregating telemetry from multiple project tenants
- Implement dashboards-as-code and alert-policy-as-code using the Terraform New Relic provider and Crossplane
- Configure New Relic alert policies routed to Opsgenie, with future integration into ServiceNow
- Manage OpenTelemetry and ADOT consumption fundamentals, including OTLP and Firehose ingestion into New Relic
- Define metric dimension strategies for large agent estates using stable, low-cardinality facets
- Operate within a Technical, Business, and Audit observability segregation model
- Enable cross-account and cross-tenant telemetry aggregation and reporting across the organization
Requirements
- 6+ years of experience in observability or SRE engineering
- Hands-on background in consolidating observability and telemetry for AI agent estates in a production agentic AI project β e.g., multi-tenant New Relic (or equivalent) aggregation of agent telemetry, cost/token attribution for GenAI workloads, or AgentCore GenAI Observability integration (generic SRE/observability experience without agent-specific telemetry context does not meet this bar)
- Proficiency in New Relic (or equivalent SaaS observability platform) operated as a multi-tenant consolidation layer, including management account and cross-account setup
- Expertise in NRQL for cross-account queries aggregating telemetry from multiple project tenants
- Skills in dashboards-as-code and alert-policy-as-code using Terraform (New Relic provider) and Crossplane
- Competency in configuring New Relic alert policies routed to Opsgenie, with awareness of ServiceNow integration
- Understanding of OpenTelemetry and ADOT consumption fundamentals, including OTLP and Firehose ingestion pipelines
- Capability to design metric dimensions for large agent estates using stable, low-cardinality facets
- Familiarity with Amazon CloudWatch and cross-tenant telemetry aggregation and reporting
- Experience working within a Technical, Business, and Audit observability segregation model
- English proficiency at an Upper-Intermediate level (B2) or higher
Nice to have
- Familiarity with AWS Bedrock AgentCore observability, including Transaction Search and GenAI Observability
- Knowledge of Amazon Web Services
- Background in Kinesis Firehose metric and log forwarding pipelines (CloudWatch β New Relic)
- Skills in cost and token attribution modeling for GenAI workloads
- Expertise in Crossplane-based provisioning of observability resources
Senior Site Reliability Engineer Β· EPAM Systems