DL
PySpark - Hyderabad
Diverse Lynx India
Location not stated
2 months ago
- PySpark
- Apache Spark
- Python
- ETL
- ELT
- SQL
- Hadoop
- Hive
- Yarn
- Parquet
- Apache Avro
- JSON
- Data Modeling
- Git
- AWS
- Azure
- GCP
- Databricks
- Airflow
- Azure Data Factory
- CI/CD
- Devops
- Delta Lake
2 months ago
Required Skills
- 4–6 years of experience in Data Engineering, Big Data Development, or related roles.
- Strong hands-on experience with PySpark and Apache Spark.
- Proficiency in Python programming.
- Experience in developing ETL/ELT workflows and data transformation processes.
- Strong SQL skills with experience in relational and analytical databases.
- Knowledge of distributed computing concepts and big data processing.
- Experience working with Hadoop ecosystem components such as HDFS, Hive, and YARN.
- Familiarity with data formats such as Parquet, Avro, ORC, JSON, and CSV.
- Understanding of data modeling and data warehousing concepts.
- Experience with Git or other version control systems.
Preferred Skills
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of Databricks or Spark-based cloud platforms.
- Familiarity with workflow orchestration tools such as Airflow, Oozie, or Azure Data Factory.
- Exposure to Kafka or other streaming technologies.
- Experience with CI/CD pipelines and DevOps practices.
- Knowledge of Delta Lake, Iceberg, or Lakehouse architectures.
Educational Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
PySpark - Hyderabad · Diverse Lynx India