Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Data & AI Reliability Engineering Consultant/Architect

EPAM Systems
  • πŸ‡ΊπŸ‡¦ Ukraine
  • Remote
  • Staff / Principal
  • 3 days ago
  • AI
  • IaC
  • New Relic
  • Datadog
  • Splunk
  • Dynatrace
  • Grafana
  • Elastic
  • OpenTelemetry
  • Incident Management
  • AWS
  • Azure
  • GCP
  • Kubernetes
  • Terraform
  • Python
  • Databricks
  • Snowflake
  • Airflow
  • RAG
  • AI/ML
  • Great Expectations
  • Unity Catalog
  • LangSmith
  • MLflow
  • Machine Learning
  • AIOps
  • ServiceNow
  • PagerDuty
  • FinOps
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seekingData & AI Reliability Engineering Consultants and Architects to drive client-facing discovery, presales and advisory engagements focused on the reliability of data platforms and AI systems. This is a consulting role: the individual leads client conversations, shapes both the solution and the proposal, and then guides the engineering team through delivery.

The consultant works across three connected domains: observability and SRE practice, data platform reliability (pipelines, quality, lineage, cost), and AI/LLM system reliability (evaluation, telemetry, guardrails).

Responsibilities

  • Drive the technical side of presales: qualify requests, facilitate client workshops, define scope, assumptions and estimates
  • Conduct discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap
  • Author the solution section of proposals and RFP responses; present and defend it to client technical and business stakeholders
  • Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths
  • Establish SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes
  • Build the business case: cost of incidents, tooling cost optimisation, expected outcomes of the change
  • Serve as the trusted advisor for client engineering leads, SDMs and directors throughout the engagement
  • Lead the first phase of delivery after a won deal, then transition to the engineering team while remaining accountable for the solution
  • Review engineers' work, set technical standards (alert-as-code, dashboards-as-code, IaC) and unblock decisions
  • Convert project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures
  • Mentor engineers moving towards consulting; participate in technical interviews
  • Represent Data & AI externally through talks, articles and vendor partnerships

Requirements

  • 7+ years in engineering, including 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership
  • Proven presales track record: leading discovery or assessment workshops, producing estimates and proposals, presenting to senior stakeholders
  • Capability to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, both in writing and live
  • English B2+ with confident spoken delivery; able to run a workshop and handle objections unsupported
  • Hands-on background in at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic
  • Knowledge of OpenTelemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction
  • Expertise in SRE practice in production: SLO/SLI, error budgets, incident management, postmortems
  • Proficiency in Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation
  • Understanding of how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming
  • Competency in data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability
  • Working understanding of LLM application architecture (RAG, agents, model gateways) and what must be measured: quality evaluation, latency, token cost, drift, guardrails
  • Background in instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client
  • Self-driven and dependable on commitments: owns deadlines for proposals and client deliverables without supervision
  • Comfort switching between several presales and one delivery engagement
  • Flexibility to use AI assistants in daily engineering and documentation work

Nice to have

  • Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP)
  • Familiarity with data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog
  • Skills in LLM observability and evaluation tooling: Langfuse, LangSmith, Arize, MLflow, OpenTelemetry GenAI conventions
  • Qualifications in AIOps and ITSM integration: ServiceNow, PagerDuty, event correlation engines
  • Proficiency in FinOps for observability and data platforms; licence and ingestion cost optimisation
  • Knowledge of AI security and governance fundamentals: guardrails, red teaming, data masking
  • Expertise in retail, finance or manufacturing domains
  • Public profile, including conference talks, articles, community leadership

Data & AI Reliability Engineering Consultant/Architect Β· EPAM Systems

Auto apply with Likeremote