ES
Senior Data Software Engineer with Databricks and Azure
EPAM Systems
๐ฐ๐ฟ Kazakhstan
Remote
Senior
2 weeks ago
- Databricks
- PySpark
- Python
- MLOps
- SAP
- HANA
- Parquet
- Delta Lake
- GitHub
- Git
- CI/CD
- GitHub Actions
- Azure
- Java
- SQL
- Scala
2 weeks ago
We are looking for aSenior Data Software Engineer to join the Data Engineering CoE, building and scaling data pipelines in Databricks with PySpark and Python to deliver data products for the client's stakeholders.
Responsibilities
- Support a team of data engineers to build pipelines used by MLOps and ML Engineers on ML modeling teams
- Develop, optimize, and maintain data transformation pipelines in Databricks (PySpark)
- Work with data stored in ADLS Gen2 and SAP HANA Data Lake, primarily in Delta/Parquet format
- Implement and maintain data quality checks, including schema validation, deduplication, enrichment, and tagging
- Communicate with stakeholders to understand business processes and model input data
- Tune performance for large-scale datasets
Requirements
- 3+ years of experience in data engineering or a related field
- Proficiency in Python, PySpark, and Databricks, including Delta Lake
- Familiarity with software version control tools such as GitHub and Git
- Experience with CI/CD frameworks such as GitHub Actions
- Knowledge of data lake technologies
- Experience working with MS Azure
- English proficiency at B2 level or higher
Nice to have
- Familiarity with at least one other programming or scripting language, such as Java, SQL, or Scala
Senior Data Software Engineer with Databricks and Azure ยท EPAM Systems