Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead Data & AI Platform Engineer

EPAM Systems
๐Ÿ‡ฆ๐Ÿ‡บ Australia
Remote
Staff / Principal
19 hours ago
  • AI
  • Machine Learning
  • CI/CD
  • Incident Response
  • Devops
  • MLOps
  • Azure
  • AWS
  • GCP
  • IaC
  • Terraform
  • Python
  • Secrets Management
  • AI/ML
  • Databricks
  • Snowflake
  • Microsoft Fabric
  • MLflow
  • RAG
  • Kubernetes
  • triage
  • GitHub Actions
  • Azure DevOps
  • GitLab CI
  • Unity Catalog
  • FinOps
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aLead Data & AI Platform Engineer to design, build, and operate the platforms that support data engineering and AI delivery across a range of client engagements. You will lead platform automation, deployment, governance, observability, and reliability for data pipelines, machine learning systems, LLM applications, and AI agents.

This is a hands-on technical leadership role. You will guide architectural decisions, establish reusable engineering practices, and work closely with data and AI teams to bring solutions into production.

Responsibilities

  • Lead the design and implementation of secure, scalable cloud infrastructure for data and AI workloads
  • Build reusable infrastructure and CI/CD patterns for data platforms, applications, models, and AI agents
  • Establish operational practices for AI solutions, including deployment, versioning, evaluation, monitoring, and rollback
  • Implement observability across pipelines and AI applications, including logs, metrics, traces, alerts, quality measures, and cost monitoring
  • Improve platform reliability through automation, incident response, performance tuning, and capacity planning
  • Define standards for cloud security, identity and access management, secrets, networking, and governance
  • Support teams deploying LLM applications and agentic systems, including their model integrations, tool calls, and external APIs
  • Use AI-assisted tools to improve infrastructure development, troubleshooting, and operational workflows, while validating their output
  • Partner with architects, engineers, and stakeholders to make platform decisions and explain technical trade-offs
  • Mentor engineers and contribute to technical standards, reusable templates, and platform roadmaps

Requirements

  • Strong hands-on experience in platform engineering, DevOps, site reliability engineering, or MLOps, with experience leading technical delivery
  • Experience with at least one major cloud platform: Azure, AWS, or Google Cloud
  • Strong infrastructure as code skills, such as Terraform
  • Experience designing CI/CD pipelines and automated deployment processes
  • Scripting or programming experience, preferably in Python
  • Experience operating data platforms, distributed workloads, or cloud-native applications
  • Strong understanding of monitoring, logging, alerting, incident response, and production troubleshooting
  • Knowledge of cloud networking, identity and access management, secrets management, and security practices
  • Practical understanding of AI/ML platform operations, including model deployment, versioning, monitoring, and lifecycle management
  • Understanding of the operational needs of LLM applications and AI agents, including tracing, evaluation, latency, reliability, and cost
  • Ability to lead technical decisions, mentor engineers, and collaborate directly with client and delivery teams

Nice to have

  • Experience with Databricks, Snowflake, Microsoft Fabric, or similar data and AI platforms
  • Experience with MLflow, model serving platforms, or model registries
  • Experience deploying and observing RAG applications or agentic systems
  • Familiarity with LLM evaluation, guardrails, and monitoring for response quality
  • Experience with Kubernetes, containers, API gateways, or service orchestration
  • Experience applying AI to operations, such as alert correlation, incident triage, or root cause analysis
  • Experience with GitHub Actions, Azure DevOps, GitLab CI, or similar tools
  • Experience with Databricks Asset Bundles, Unity Catalog, or equivalent capabilities
  • FinOps experience, including workload optimisation and cloud cost attribution
  • Consulting experience, including platform assessments, solution design, and technical estimation

Lead Data & AI Platform Engineer ยท EPAM Systems

Auto apply with Likeremote