ES
Senior Data Software Engineer/ Databricks, Apache Spark, PySpark
EPAM Systems
π¦π² Armenia | π°π¬ Kyrgyzstan | π°πΏ Kazakhstan | πΊπΈ United States | πΊπΏ Uzbekistan
Remote
Senior
18 hours ago
- PySpark
- SQL
- Unit Testing
- Apache Spark
- Databricks
- pytest
- ETL
- Git
- Azure
- AWS
- GCP
- CI/CD
- Devops
- Delta Lake
- MLflow
18 hours ago
We are seeking a skilledSenior Data Software Engineerwith strong expertise inPySpark,SQL, andunit testing to join our data engineering team. The ideal candidate will have hands-on experience withApache Spark, preferably within theDatabricks environment, and will be responsible for building scalable data pipelines, optimizing data workflows, and ensuring code quality through rigorous testing practices.
Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark (Databricks preferred)
- Write efficient and optimized PySpark code for data transformation and processing
- Develop and execute complex SQL queries for data extraction, validation, and reporting
- Implement unit tests using pytest to ensure code reliability and maintainability
- Collaborate with data scientists, analysts, and other engineers to deliver high-quality data solutions
- Monitor and troubleshoot data workflows and performance issues
- Document technical designs, processes, and best practices
Requirements
- 3+ years of experience in Data Software Engineering
- Proven experience with Apache Spark, ideally in a Databricks environment
- Proficiency in PySpark and SQL
- Background in unit testing frameworks, especially pytest
- Understanding of data engineering principles and ETL processes
- Familiarity with version control systems (e.g., Git)
- Ability to work independently and in a collaborative team setting
- Excellent problem-solving and communication skills
- Proficiency in English at an Upper-Intermediate level (B2) or higher
Nice to have
- Experience with cloud platforms (e.g., Azure, AWS, GCP)
- Knowledge of CI/CD pipelines and DevOps practices
- Familiarity with Delta Lake, MLflow, or other Databricks-native tools
Senior Data Software Engineer/ Databricks, Apache Spark, PySpark Β· EPAM Systems