Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Site Reliability Engineer

EPAM Systems
πŸ‡ΊπŸ‡¦ Ukraine
Remote
Senior
22 hours ago
  • AI
  • LangGraph
  • AWS
  • CI/CD
  • New Relic
  • Terraform
  • Opsgenie
  • ServiceNow
  • OpenTelemetry
  • Machine Learning
  • CloudWatch
  • AWS Bedrock
  • Kinesis
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for aSenior Site Reliability Engineerto join our team in building an Enterprise Agent Development Platform β€” a production-grade, cloud-native ecosystem that enables engineering teams to define, orchestrate, deploy, and observe AI agents at scale. The platform standardizes agent development across the organization using LangGraph and Strands Agents on AWS AgentCore Runtime. This initiative spans agent framework design, runtime architecture, marketplace integration, CI/CD automation, and enterprise-grade observability β€” reducing agent development from months to days while enforcing consistent security, quality, and governance standards.

As a Senior Observability Engineer focused on New Relic Consolidation, you will own the multi-tenant New Relic consolidation layer, aggregating telemetry across the agent estate.

Responsibilities

  • Own the multi-tenant New Relic consolidation layer aggregating telemetry across the entire agent estate
  • Administer New Relic at scale, including management account setup and cross-account configuration
  • Design and maintain NRQL cross-account queries aggregating telemetry from multiple project tenants
  • Implement dashboards-as-code and alert-policy-as-code using the Terraform New Relic provider and Crossplane
  • Configure New Relic alert policies routed to Opsgenie, with future integration into ServiceNow
  • Manage OpenTelemetry and ADOT consumption fundamentals, including OTLP and Firehose ingestion into New Relic
  • Define metric dimension strategies for large agent estates using stable, low-cardinality facets
  • Operate within a Technical, Business, and Audit observability segregation model
  • Enable cross-account and cross-tenant telemetry aggregation and reporting across the organization

Requirements

  • 6+ years of experience in observability or SRE engineering
  • Hands-on background in consolidating observability and telemetry for AI agent estates in a production agentic AI project β€” e.g., multi-tenant New Relic (or equivalent) aggregation of agent telemetry, cost/token attribution for GenAI workloads, or AgentCore GenAI Observability integration (generic SRE/observability experience without agent-specific telemetry context does not meet this bar)
  • Proficiency in New Relic (or equivalent SaaS observability platform) operated as a multi-tenant consolidation layer, including management account and cross-account setup
  • Expertise in NRQL for cross-account queries aggregating telemetry from multiple project tenants
  • Skills in dashboards-as-code and alert-policy-as-code using Terraform (New Relic provider) and Crossplane
  • Competency in configuring New Relic alert policies routed to Opsgenie, with awareness of ServiceNow integration
  • Understanding of OpenTelemetry and ADOT consumption fundamentals, including OTLP and Firehose ingestion pipelines
  • Capability to design metric dimensions for large agent estates using stable, low-cardinality facets
  • Familiarity with Amazon CloudWatch and cross-tenant telemetry aggregation and reporting
  • Experience working within a Technical, Business, and Audit observability segregation model
  • English proficiency at an Upper-Intermediate level (B2) or higher

Nice to have

  • Familiarity with AWS Bedrock AgentCore observability, including Transaction Search and GenAI Observability
  • Knowledge of Amazon Web Services
  • Background in Kinesis Firehose metric and log forwarding pipelines (CloudWatch β†’ New Relic)
  • Skills in cost and token attribution modeling for GenAI workloads
  • Expertise in Crossplane-based provisioning of observability resources

Senior Site Reliability Engineer Β· EPAM Systems

Auto apply with Likeremote