MLOps / AI Operations Engineer
- AI
- CI/CD
- Devops
- IaC
- Machine Learning
- GitHub Actions
- Azure DevOps
- GitLab CI
- Terraform
- Bicep
- Docker
- Kubernetes
- MLflow
- OpenTelemetry
- LangSmith
- Azure Monitor
- Prometheus
- Grafana
- Secrets Management
- Incident Response
- Risk Management
Job Description & Summary
The opportunity
Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management.
What you will be doing
·       Build CI/CD pipelines for AI services, prompts, agent configurations, infrastructure and evaluation assets.
·       Automate environment provisioning, testing, deployment, rollback and release evidence.
·       Implement tracing, logging, model and agent monitoring, alerts and operational dashboards.
·       Operationalize evaluation thresholds, incident handling and continuous-improvement loops.
·       Monitor latency, capacity, token usage, infrastructure consumption and cost drivers.
·       Define runbooks, service ownership and production support handover.
What we need from you
·       4+ years in DevOps, platform engineering, ML engineering, SRE or cloud operations.
·       Strong automation, containers, cloud services, observability and Infrastructure as Code capability.
·       Experience deploying or operating ML, generative AI or distributed application workloads.
·       Understanding of release controls, reliability, security and cost optimization.
Relevant AI technologies and tooling
·       Hands-on experience with GitHub Actions, Azure DevOps, GitLab CI or equivalent, plus Infrastructure as Code using Terraform, Bicep or comparable tooling.
·       Strong container and orchestration capability using Docker and Kubernetes, together with experience deploying AI or agent services across cloud and hybrid environments.
·       Experience operating model and prompt assets, agent configurations, evaluation datasets and release evidence using MLflow, platform-native registries or equivalent lifecycle tooling.
·       Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith, MLflow, Langfuse, Azure Monitor, Prometheus or Grafana.
·       Ability to monitor model and agent quality, tool failures, retrieval performance, latency, token usage, cost, capacity and workflow-level service indicators.
·       Experience with progressive delivery, rollback, secrets management, vulnerability scanning, incident response and reliability practices for non-deterministic AI systems.
Measures of success
·       Deployment frequency and success rate
·       Mean time to detect and restore
·       Evaluation and monitoring coverage
·       Service reliability and latency
·       Cost and consumption transparency
Key interfaces
·       Other members of the AI Transformation & Agentic Systems Practice
·       PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
·       Client business owners, product owners, technology teams and operational users
·       Technology alliance and implementation partners where relevant
Contribution to the practice
·       Support proposals, client workshops and market development appropriate to seniority.
·       Contribute reusable methods, patterns, code, assets and lessons learned.
·       Coach colleagues and participate in the capability’s continuous learning agenda.
·       Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-HybridÂ
MLOps / AI Operations Engineer · wd3:pwc:Global_Experienced_Careers