EI
Lead Azure Data Engineer
eTeam Inc.
🇺🇸 United States
Hybrid
Staff / Principal
1 week ago
- Azure Data Factory
- Azure Databricks
- Scala
- Python
- PySpark
- SQL
- Azure Synapse
- JSON
- Databricks
- pytest
- CI/CD
- Azure DevOps
- GitHub Actions
- IaC
- Terraform
- Bicep
- Key Vault
- Delta Lake
- Unity Catalog
- Azure
- Azure SQL
- Git
- Data Modeling
- Parquet
- Airflow
- MLflow
- Devops
1 week ago
Work Location & Reporting Address: Bellevue, WA 98006 (2 days a week in Office)
Job Details:
Minimum years of experience required: 8
Certification needed: NA
Must Have Skills: Azure Data factory, Azure Databricks
Nice to Have Skills: Scala, Python (PySpark), SQL
Key Responsibilities
•Design, build, and maintain PySpark/SQL pipelines in Azure Databricks for batch and streaming data.
•Develop robust ingestion from Azure Data Lake Storage (ADLS Gen2), Azure Synapse/SQL, Event Hub, Kafka, and REST/JSON sources.
•Optimize Spark jobs (partitioning, caching, broadcast joins, AQE) for performance and cost.
•Implement monitoring and alerting (cluster/job metrics, driver/executor logs).
•Use Databricks Repos, notebooks, and modular PySpark projects with unit tests (pytest).
•Build CI/CD pipelines (e.g., Azure DevOps, GitHub Actions) for jobs, notebooks, and infrastructure-as-code (Terraform/ARM/Bicep).
•Manage environments (dev/test/prod), secrets/Key Vault, and configuration promotion.
Required Qualifications (Intermediate Level)
•6 years in data engineering; 4 years hands-on with Azure Databricks and Spark.
•Strong PySpark and SQL skills: DataFrames, joins, window functions, UDFs, incremental loads.
•Practical experience with Delta Lake, Unity Catalog, and Databricks Jobs/Workflows.
•Familiarity with Azure services: ADLS Gen2, Azure Key Vault, Event Hub, Azure SQL/Synapse.
•Version control (Git) and CI/CD experience; basic testing practices (pytest).
•Ability to optimize Spark jobs and troubleshoot: skew, shuffle, OOM, driver/executor tuning.
•Solid understanding of data modeling (star schema, medallion/lakehouse), partitioning, and file formats (Parquet/JSON).
•Airflow, Azure Data Factory orchestration.
•Terraform for Databricks & Azure resources.
•Basic Scala and/or SQL Warehouses (Databricks SQL) for BI.
Education
•Bachelor’s/Master’s in Computer Science, Engineering, or related field (or equivalent experience).
Certifications (Optional but Valued)
•Databricks: Data Engineer Associate/Professional
•Microsoft Azure: DP-203 (Data Engineering on Microsoft Azure), AZ-900 (Fundamentals)
Tools & Tech Stack (Typical)
•Languages: Python (PySpark), SQL
•Databricks: Notebooks, Jobs/Workflows, Repos, Unity Catalog, Delta Lake, MLflow
•Azure: ADLS Gen2, Key Vault, Event Hub, Synapse/SQL, Monitor/Log Analytics
•DevOps: Git, Azure DevOps/GitHub Actions, Terraform/Bicep
Lead Azure Data Engineer · eTeam Inc.