Databrick engineer
- Python
- SQL
- Scala
- Apache Spark
- PySpark
- Azure Databricks
- Azure Synapse
- Delta Lake
- Parquet
- ETL
- ELT
- Git
- Azure DevOps
- CI/CD
Databrick engineer
• Languages: Python, SQL, Scala
• Processing: Apache Spark, PySpark, Spark SQL
• Platform: Azure Databricks (notebooks, Jobs/Workflows, clusters), Azure Synapse
• Storage/format: Delta Lake, Parquet, ADLS Gen2
• Techniques: ETL/ELT, MERGE/upsert, partitioning, Spark performance tuning
• Tooling: Git, Azure DevOps / CI/CD
Techniques (Data Engineering Practices)
• ETL / ELT
• Extract data from source systems
• Transform/clean/reshape into usable form
• Load into target tables (ELT = transform after loading, inside the platform)
• MERGE / Upsert
• Update existing rows when records already exist
• Insert new rows when they don't
• Handled in a single atomic operation (core Delta Lake pattern)
• Partitioning
• Split large tables by a column (e.g., date, source type)
• Enables partition pruning — scan only relevant data
• Improves query speed and reduces cost
• Spark Performance Tuning
• Handle data skew and manage shuffles
• Caching / persistence of reused datasets
• Right-size partitions; use broadcast joins
• Cluster sizing and resource optimization
Databrick engineer · Info Way Solutions LLC