Developer
- Python
- PySpark
- Apache Spark
- Hadoop
- Hive
- SQL
- REST API
- JSON
- XML
- Machine Learning
- ETL
- AWS
- Azure
- GCP
Must-Have**
Strong proficiency inPython programming.
Hands-on experience withPySpark andApache Spark.
Knowledge ofBig Data technologies (Hadoop, Hive, Kafka, etc.).
Experience withSQL and relational/non-relational databases.
Familiarity withdistributed computing andparallel processing.
Understandingdata engineering best practices.
Experience withREST APIs,JSON/XML, and data serialization.
Exposure tocloud computing environments.
5+ years of experience in Python and PySpark development.
Experience withdata warehousing anddata lakes.
Knowledge ofmachine learning libraries (e.g., MLlib) is a plus.
Strong problem-solving and debugging skills.
Excellent communication and collaboration abilities.
Develop and maintain scalable data pipelines usingPython andPySpark.
Design and implementETL (Extract, Transform, Load) processes.
Optimize and troubleshoot existing PySpark applications for performance.
Collaborate with cross-functional teams to understand data requirements.
Write clean, efficient, and well-documented code.
Conduct code reviews and participate in design discussions.
Ensure data integrity and quality across the data lifecycle.
Integrate with cloud platforms likeAWS,Azure, orGCP.
Implement data storage solutions and manage large-scale datasets.
Developer ยท GSB Solutions