Description:
• a bachelor's degree or higher in Computer Science/ Engineering/ Science
• data engineer with 8 to 10 years of experience, having 5+ experience in Azure cloud native services.
• Experience in implementing and Designing Big Data solutions in Azure Databricks for Batch and Real Time Analytics uses cases.
• Experience in core concepts like Lakehouse, Delta Lake, Data Mesh, Data Virtualization, Dimensional Modelling techniques, Data Aggregation techniques.
• Experience in programming skills in languages such as Python/Pyspark and frameworks and libraries based on them with a focus on data engineering.
• Extensive experience designing, building, and operating robust distributed data platforms (e.g., Spark, Kafka, Flink, HBase) and handling data at the petabyte scale.
• Experience in data engineering on Apache Spark and populate downstream batch data stores (such as data mart) for BI use cases & to generate downstream feeds (ie., flat & wide tables or compressed files) for Data Science use cases
• Experience in Real-time service integration to process business events off Kafka and persist in ADLS Gen 2 using Databricks.
• Experience in Near real-time stream processing to derive features for ML model inference.
• Experience in Hadoop based implementations, Python, Spark, SQL, Unix, and Hive
• Extensive experience and knowledge of a variety of data technologies and frameworks such as Delta Lake, Datahub, Apache Spark, Azure Databricks, Databricks Genie
• Strong familiarity with the concepts of RDBMS, Data Lake, Data Warehouse, Data Lakehouse, and Medallion Architecture.
• Knowledge of at least one workflow orchestration tools such as Apache Airflow, Luigi, Oozie, AWS Glue or similar frameworks
• Experience in Azure data bricks performance optimization techniques (Liquid clustering, vacuums, Z-ordering etc)
• Experience in Azure Services like Azure Data Factory, Azure Data Bricks (Scala/Python), ADLS Gen2, Azure SQL Database, SQL/T-SQL
• Experience working with the DevOps teams to drive operationally excellent infrastructure.
• Knowledge on Microsoft Fabric and its framework
• Basic understanding of data quality dimensions like Consistency, Completeness, Accuracy, Lineage etc.
• Understanding on different monitoring and logging solutions on Azure
• Good to have experience on Data Visualization Tools like Power BI and Tableau
• Good to have experience of Apache Flink/Apache Kafka and the ELK stack (Elasticsearch, Logstash & Kibana)
• great communicator, you know how to speak with people at all levels.
• team player, working with a globally distributed team. |
work location:
|
|
|
1. Azure Databricks, Azure Data Factory (ADF), ADLS Gen2, Azure SQL
2. Python, PySpark, Apache Spark, SQL/T-SQL
3. Lakehouse, Delta Lake, Data Mesh, Data Virtualization
4. Dimensional Modeling & Data Aggregation techniques
5. Kafka and Real-time/Near Real-time Data Processing
6. Hadoop Ecosystem, Hive, Unix
7. Data Lake, Data Warehouse, Medallion Architecture
8. Workflow Orchestration (Airflow/Luigi/Oozie/AWS Glue)
9. Databricks Performance Tuning (Liquid Clustering, Z-Ordering, Vacuum)
10. Data Quality concepts (Accuracy, Completeness, Consistency, Lineage)
11. Azure Monitoring & Logging
12. DevOps collaboration experience
13. Strong communication and teamwork skills |