ES
Data & AI Reliability Engineering Consultant/Architect
EPAM Systems
- πΊπ¦ Ukraine
- Remote
- Staff / Principal
- 3 days ago
- AI
- IaC
- New Relic
- Datadog
- Splunk
- Dynatrace
- Grafana
- Elastic
- OpenTelemetry
- Incident Management
- AWS
- Azure
- GCP
- Kubernetes
- Terraform
- Python
- Databricks
- Snowflake
- Airflow
- RAG
- AI/ML
- Great Expectations
- Unity Catalog
- LangSmith
- MLflow
- Machine Learning
- AIOps
- ServiceNow
- PagerDuty
- FinOps
3 days ago
We are seekingData & AI Reliability Engineering Consultants and Architects to drive client-facing discovery, presales and advisory engagements focused on the reliability of data platforms and AI systems. This is a consulting role: the individual leads client conversations, shapes both the solution and the proposal, and then guides the engineering team through delivery.
The consultant works across three connected domains: observability and SRE practice, data platform reliability (pipelines, quality, lineage, cost), and AI/LLM system reliability (evaluation, telemetry, guardrails).
Responsibilities
- Drive the technical side of presales: qualify requests, facilitate client workshops, define scope, assumptions and estimates
- Conduct discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap
- Author the solution section of proposals and RFP responses; present and defend it to client technical and business stakeholders
- Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths
- Establish SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes
- Build the business case: cost of incidents, tooling cost optimisation, expected outcomes of the change
- Serve as the trusted advisor for client engineering leads, SDMs and directors throughout the engagement
- Lead the first phase of delivery after a won deal, then transition to the engineering team while remaining accountable for the solution
- Review engineers' work, set technical standards (alert-as-code, dashboards-as-code, IaC) and unblock decisions
- Convert project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures
- Mentor engineers moving towards consulting; participate in technical interviews
- Represent Data & AI externally through talks, articles and vendor partnerships
Requirements
- 7+ years in engineering, including 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership
- Proven presales track record: leading discovery or assessment workshops, producing estimates and proposals, presenting to senior stakeholders
- Capability to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, both in writing and live
- English B2+ with confident spoken delivery; able to run a workshop and handle objections unsupported
- Hands-on background in at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic
- Knowledge of OpenTelemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction
- Expertise in SRE practice in production: SLO/SLI, error budgets, incident management, postmortems
- Proficiency in Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation
- Understanding of how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming
- Competency in data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability
- Working understanding of LLM application architecture (RAG, agents, model gateways) and what must be measured: quality evaluation, latency, token cost, drift, guardrails
- Background in instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client
- Self-driven and dependable on commitments: owns deadlines for proposals and client deliverables without supervision
- Comfort switching between several presales and one delivery engagement
- Flexibility to use AI assistants in daily engineering and documentation work
Nice to have
- Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP)
- Familiarity with data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog
- Skills in LLM observability and evaluation tooling: Langfuse, LangSmith, Arize, MLflow, OpenTelemetry GenAI conventions
- Qualifications in AIOps and ITSM integration: ServiceNow, PagerDuty, event correlation engines
- Proficiency in FinOps for observability and data platforms; licence and ingestion cost optimisation
- Knowledge of AI security and governance fundamentals: guardrails, red teaming, data masking
- Expertise in retail, finance or manufacturing domains
- Public profile, including conference talks, articles, community leadership
Data & AI Reliability Engineering Consultant/Architect Β· EPAM Systems