
Data Engineer
Milestone Technologies
- ๐ฎ๐ณ India
- 17 hours ago
- SQL
- Python
- Scala
- Apache Spark
- AWS
- EMR
- Step Functions
- Apache Iceberg
- dbt
- Airflow
- AWS Glue
- AI
17 hours ago
Data Engineer โ Responsibilities (4โ7 Years Exp)
- Develop and maintain scalabledata ingestion and processing pipelines using SQL, Spark, and Python/Scala.
- Develop and optimizeApache Spark batch and streaming jobs for large-scale data processing.
- Design and implementdata models includingDimension and Fact tables, with clearly defined grain, relationships, and measures.
- ImplementSlowly Changing Dimensions (SCD) and appropriate incremental processing strategies.
- Work withAWS data engineering and serverless services such as Glue, S3, Lambda, EMR, and Step Functions.
- Develop and maintain reliable, scalable data pipelines with appropriateerror handling, logging, monitoring, and data quality checks.
- Apply data engineering fundamentals includingpartitioning, incremental processing, schema evolution, data quality, and pipeline reliability.
- Work withApache Iceberg, including table design, partitioning, schema evolution, and incremental processing.
- Develop data transformation workflows usingDBT and orchestrate pipelines usingAirflow.
- Troubleshoot and optimize data pipelines forperformance, scalability, and cost efficiency.
- Participate incode reviews, testing, debugging, and continuous improvement of data engineering solutions.
Required Skills
- 4โ7 years of relevant Data Engineering experience.
- Strong SQL skills โ SQL is a core requirement.
- Strong hands-on experience withApache Spark andPython or Scala.
- Hands-on experience withAWS data engineering and serverless services, particularly AWS Glue and S3.
- Good understanding ofData Modelling, including:
- Dimension and Fact table design
- Defining appropriateFact grain
- Dimension-to-Fact relationships
- ImplementingSCD Type 1 / Type 2 and other appropriate SCD patterns
- Strong understanding of data engineering fundamentals includingpartitioning, incremental processing, schema evolution, data quality, and pipeline reliability.
- Good understanding ofdistributed data processing and Spark performance optimization.
Additional Responsibilities / Good to Have
- Experience withApache Iceberg and lakehouse architectures.
- Experience withDBT for data transformation and modelling.
- Experience withAirflow for workflow orchestration.
- Kafka / Kafka Streaming experience is a good to have.
- Exposure toAI/LLM tools for development, debugging, testing, and documentation.
Data Engineer ยท Milestone Technologies