Senior Data Engineer/ Databricks, PySpark, Python
- Databricks
- Azure Databricks
- PySpark
- Azure
- Python
- CI/CD
- GitHub Actions
- IoT
We are looking for an experiencedSenior Data Engineer / Databricks Engineer to join a project focused on collecting, processing, and analyzing vehicle telemetry data for a major European automotive manufacturer.
The platform operates at a significant scale, processing tens of terabytes of telemetry data in Azure Databricks and running approximately 200 Databricks jobs. The solution is currently in the production maintenance and evolution phase and is deployed across multiple regions, including Europe, North America, and China, with an additional deployment in China planned in the near future.
A key focus of the role will be ensuring that the system can scale reliably while keeping infrastructure and operational costs predictable.
Responsibilities
- Monitor the health, stability, and performance of the production data platform
- Investigate and resolve performance issues and production incidents
- Analyze Databricks workloads and identify performance bottlenecks
- Optimize PySpark workloads and Databricks jobs for performance, reliability, and cost efficiency
- Support and improve approximately 200 existing Databricks jobs and associated data pipelines
- Prepare the platform for increasing data volumes and workloads as additional vehicle platforms are onboarded
- Design and implement scaling strategies with a strong focus on predictable infrastructure and operational costs
- Track and optimize Databricks and Azure resource consumption
- Improve platform observability, monitoring, alerting, and operational processes
- Ensure consistency across multiple regional deployments in Europe, North America, and China
- Identify opportunities for technical improvements, automation, and reduction of operational overhead
Requirements
- 3+ years of hands-on experience in data engineering
- Expertise in Azure Databricks
- Proficiency in PySpark and distributed data processing
- Advanced skills in Python development
- Understanding of performance optimization for large-scale Spark workloads
- Background in operating and troubleshooting production data platforms
- Knowledge of monitoring, observability, and capacity planning combined with cloud cost optimization
- Ability to analyze existing systems, identify bottlenecks, and propose pragmatic improvements
- English proficiency at an Upper-Intermediate level (B2) or higher
Nice to have
- Familiarity with Microsoft Azure services and cloud infrastructure
- Experience building and maintaining CI/CD pipelines using GitHub Actions
- Exposure to large-scale telemetry, IoT, or automotive data
- Background in managing Databricks environments with a large number of scheduled jobs and pipelines
- Showcase of multi-region or geographically distributed cloud deployments
Senior Data Engineer/ Databricks, PySpark, Python Β· EPAM Systems