Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
PL

Data Engineer with data science exp

PeopleNTech LLC
๐Ÿ‡บ๐Ÿ‡ธ United States
On-site
5 months ago
  • PySpark
  • Excel
  • SAS
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Role : Data Engineer with data science exp
Location : Scottsdale AZ (Onsite)
Rate : $75/hr.
Indent :
SF_OP_201150-1-1

We are looking for a skilledData Engineer with strong PySpark experience to work on large-scale data processing and analytics initiatives. The ideal candidate will have hands-on experience working withlarge datasets, complex joins, and performance optimization, along with the ability to applybasic analytical thinking and deliverclear, stakeholder-ready outputs.

Key Responsibilities
Data Engineering & Development
  • Design, develop, and maintain scalable data pipelines usingPySpark.
  • Writeefficient and optimized PySpark code to process and transformlarge-scale datasets.
  • Handlejoins across multiple large databases, ensuring performance, accuracy, and scalability.
  • Optimize Spark jobs tominimize runtime, memory usage, and compute cost.
  • Work with structured and semi-structured data from multiple sources.
Data Preparation & Analysis Support
  • Build and curatetraining and analytical datasets by joining and transforming multiple data sources.
  • Applybasic analytical skills to understand data patterns, anomalies, and business relevance.
  • Performdata validation and quality checks, including:
    • Record counts and reconciliation
    • Duplicate detection
    • Null and outlier checks
    • Schema and data-type validation
  • Ensure datasets areanalysis-ready and trustworthy.
Stakeholder Interaction & Reporting
  • Understand business objectives and translate them into data requirements.
  • Ask the right questions to determine:
    • Level of aggregation required
    • Metrics definitions
    • Data freshness and accuracy expectations
    • Preferred output and reporting formats
  • Present results and insights clearly to stakeholders.
  • Createreports and summaries using Excel for business users and leadership.

Expected Technical Approach (Problem-Solving Mindset)
Candidates are expected to demonstrate the ability to:
  • Approach complex data projects methodically, starting with:
    • Understanding business objectives
    • Reviewing source data structure and volume
    • Designing efficient join strategies
  • Choose the right join types, partitioning strategies, and caching techniques.
  • Validate data at every stage of the pipeline.
  • Balance technical accuracy with business usability when presenting results.

Core Skill Sets (Must-Have)
  • Strong hands-on experience with PySpark
  • Extensive experience working with large datasets
  • Proven expertise injoining large databases efficiently
  • Ability to writehigh-performance, optimized code
  • Basic analytical skills to interpret and validate data
  • Reporting skills using Excel

Good to Have Skills
  • Experience inmodel development or supporting analytics/modeling teams
  • SAS experience
  • Exposure toCloudera or similar big data platforms
  • Understanding of data warehousing and analytics workflows

Soft Skills & Competencies
  • Strong problem-solving and logical thinking

DISCLAIMER
==========
This e-mail may contain privileged and confidential information which is the property of Persistent Systems Ltd. It is intended only for the use of the individual or entity to which it is addressed. If you are not the intended recipient, you are not authorized to read, retain, copy, print, distribute or use this message. If you have received this communication in error, please notify the sender and delete all copies of this message. Persistent Systems Ltd. does not accept any liability for virus infected mails.

Data Engineer with data science exp ยท PeopleNTech LLC

Auto apply with Likeremote