Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
DL

Pyspark_tcs

Diverse Lynx India
Location not stated
Senior
2 months ago
  • PySpark
  • ETL
  • ELT
  • Apache Spark
  • Python
  • SQL
  • Hadoop
  • Hive
  • Yarn
  • Parquet
  • Apache Avro
  • JSON
  • Data Modeling
  • Git
  • AWS
  • Azure
  • GCP
  • Databricks
  • Airflow
  • Azure Data Factory
  • CI/CD
  • Devops
  • Delta Lake
  • Agile
  • Scrum
  • Machine Learning
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Role Descriptions: Sr. Spark Developers
Experience Required: 6-8
Skills: Digital: PySpark
Location: ~HYDERABAD~

Job Summary

We are looking for a Junior Spark Developer with strong expertise in PySpark to develop, optimize, and maintain large-scale data processing applications. The ideal candidate should have hands-on experience in building ETL pipelines, processing big data workloads, and working with distributed computing frameworks. The role involves collaborating with data engineers, data architects, and business stakeholders to deliver scalable and efficient data solutions.

Key Responsibilities
- Design, develop, and maintain data processing applications using PySpark.
- Build and optimize ETL/ELT pipelines for large-scale data ingestion and transformation.
- Work with structured, semi-structured, and unstructured data from multiple sources.
- Develop Spark jobs for batch and near real-time data processing.
- Optimize Spark applications for performance, scalability, and reliability.
- Collaborate with Data Engineering and Analytics teams to understand business requirements and implement data solutions.
- Perform data validation, data quality checks, and troubleshooting of data pipelines.
- Monitor and support production data workflows and resolve issues promptly.
- Write reusable, maintainable, and efficient code following best practices.
- Participate in code reviews, testing, and deployment activities.
- Create and maintain technical documentation for developed solutions.

Required Skills
- 4–6 years of experience in Data Engineering, Big Data Development, or related roles.
- Strong hands-on experience with PySpark and Apache Spark.
- Proficiency in Python programming.
- Experience in developing ETL/ELT workflows and data transformation processes.
- Strong SQL skills with experience in relational and analytical databases.
- Knowledge of distributed computing concepts and big data processing.
- Experience working with Hadoop ecosystem components such as HDFS, Hive, and YARN.
- Familiarity with data formats such as Parquet, Avro, ORC, JSON, and CSV.
- Understanding of data modeling and data warehousing concepts.
- Experience with Git or other version control systems.

Preferred Skills
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of Databricks or Spark-based cloud platforms.
- Familiarity with workflow orchestration tools such as Airflow, Oozie, or Azure Data Factory.
- Exposure to Kafka or other streaming technologies.
- Experience with CI/CD pipelines and DevOps practices.
- Knowledge of Delta Lake, Iceberg, or Lakehouse architectures.

Educational Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.

Key Competencies
- Strong Analytical and Problem-Solving Skills
- Data Processing and Optimization
- Attention to Detail
- Team Collaboration
- Communication Skills
- Continuous Learning and Adaptability

Nice to Have
- Experience in enterprise-scale data engineering projects.
- Exposure to Agile/Scrum development methodologies.
- Spark or cloud platform certifications.
- Basic understanding of machine learning data pipelines and analytics workflows.

Pyspark_tcs · Diverse Lynx India

Auto apply with Likeremote