ES
Lead Data & AI Engineer
EPAM Systems
🇳🇱 Netherlands
Hybrid
Staff / Principal
19 hours ago
- AI
- ETL
- ELT
- AI/ML
- Azure Data Factory
- Databricks
- MLOps
- Machine Learning
- SQL
- Python
- PySpark
- Delta Lake
- CI/CD
- GitHub Actions
- Azure DevOps
- RAG
- Foundry
- GitHub
- Azure
- dbt
19 hours ago
We’re looking for aLead Data & AI Engineer to join our team in the Netherlands in a hybrid working mode. In this role, you will design and deliver advanced data solutions and AI-assisted capabilities on modern cloud platforms. You’ll build scalable data pipelines, optimize data quality and governance and enable ML feature engineering for analytics and intelligent applications. The role requires hands-on technical expertise in data engineering and AI integration while driving reliability, performance and security for enterprise-scale systems.
Responsibilities
- Design and implement robust data architectures using cloud-native technologies
- Build large-scale ETL/ELT workflows to process heterogeneous datasets
- Create data models optimized for AI/ML pipelines and advanced analytics
- Develop streaming and batch pipelines leveraging tools like Azure Data Factory and Databricks
- Operationalize ML solutions, integrating feature stores, model registries and inference endpoints
- Collaborate with data scientists to deploy and monitor AI models using enterprise MLOps frameworks
- Integrate GenAI-assisted development methods into data workflows for automation and efficiency
- Ensure data governance, lineage, cataloging and quality frameworks across platforms
- Optimize platform performance, manage compute cost efficiency and enable observability for critical workloads
- Contribute to best practices, mentoring engineering teams in advanced data and AI engineering
Requirements
- 5+ years working in data engineering, with proven architecture and leadership experience
- Advanced proficiency in SQL for performance tuning at large scale
- Strong programming skills in Python, with applied experience in data engineering workflows
- Expertise in PySpark for distributed data processing
- Hands-on experience with Databricks, including Delta Lake and performance optimization
- Knowledge of Azure Data Factory for orchestration of pipelines (ETL/ELT)
- Familiarity with AI/ML pipeline development, including integration of models into production
- Demonstrated exposure to Gen AI-assisted development workflows for accelerating data engineering tasks
- Strong knowledge of CI/CD for data and ML pipelines (GitHub Actions, Azure DevOps) and experience with large-scale environments
- English proficiency at B2 level (Upper-Intermediate) or higher
Nice to have
- Prompt Engineering knowledge and experience building RAG workflows
- Familiarity with Microsoft Foundry platforms
- Version control and CI/CD experience with GitHub
- Understanding of ETL/ELT optimization patterns beyond Azure stack
- Knowledge of data mesh or lakehouse architectural concepts
- Hands-on exposure to dbt (data build tool), MKDocs and similar developer productivity tooling
Lead Data & AI Engineer · EPAM Systems