IW
Data Engineer
Info Way Solutions LLC
🇦🇿 Azerbaijan
On-site
4 weeks ago
- Python
- SQL
- Scala
- Apache Spark
- PySpark
- Azure Databricks
- Azure Synapse
- Delta Lake
- Parquet
- ETL
- ELT
- Git
- Azure DevOps
- CI/CD
4 weeks ago
Data Engineer
Phoenix, AZ// Seattle, WA
Data:
• Languages: Python, SQL, Scala
• Processing: Apache Spark, PySpark, Spark SQL
• Platform: Azure Databricks (notebooks, Jobs/Workflows, clusters), Azure Synapse
• Storage/format: Delta Lake, Parquet, ADLS Gen2
• Techniques: ETL/ELT, MERGE/upsert, partitioning, Spark performance tuning
• Tooling: Git, Azure DevOps / CI/CD
Techniques (Data Engineering Practices)
• ETL / ELT
• Extract data from source systems
• Transform/clean/reshape into usable form
• Load into target tables (ELT = transform after loading, inside the platform)
• MERGE / Upsert
• Update existing rows when records already exist
• Insert new rows when they don't
• Handled in a single atomic operation (core Delta Lake pattern)
• Partitioning
• Split large tables by a column (e.g., date, source type)
• Enables partition pruning — scan only relevant data
• Improves query speed and reduces cost
• Spark Performance Tuning
• Handle data skew and manage shuffles
• Caching / persistence of reused datasets
• Right-size partitions; use broadcast joins
• Cluster sizing and resource optimization
Data Engineer · Info Way Solutions LLC