TI
Senior Data Engineer
Tranzeal Inc.
🇮🇳 India
Hybrid
Senior
1 week ago
- Apache Spark
- AWS Glue
- Athena
- Power BI
- SQL
- AI
- Devops
- EDI
- AWS
- ETL
- ELT
- PySpark
- DAX
- Power Query
- Data Modeling
- Claude Code
- Cursor
- HIPAA
- X12
- HL7
- FHIR
- Redshift
- EMR
- Health insurance
1 week ago
Location: Hyderabad, India (Hybrid - 3 Days WFO)
Experience: 5–7 years
CTC : 21.5 LPA
Notice period : Immediate to 30 Days Max
Tech Stack: Apache Spark (batch and streaming), AWS Glue, S3, Athena, Power BI, SQL, AI-Assisted Development
Role Overview
We are looking for a Senior Data Engineer to build and run the data and reporting layer of our health plan technology platform — batch and streaming pipelines on AWS Glue and Spark, and the Power BI models and reports built on top of them. The role sits within the Data/BI & Integration team alongside software engineering, DevOps and Edifecs EDI, and consumes event and file feeds from the platform's integration layer as well as core AWS sources.
Key Responsibilities
• Design, build, and maintain batch ETL/ELT pipelines using AWS Glue and PySpark across S3 and Athena
• Build and maintain streaming pipelines in Spark for near-real-time ingestion of claims, eligibility, and enrollment events
• Own the Glue Data Catalog — crawlers, table definitions, job orchestration, and scheduling
• Build and maintain Power BI semantic models and reports — star-schema datasets, DAX measures, and Power Query transformations
• Write and optimize SQL supporting pipelines, reporting datasets, and downstream applications
• Design and maintain dimensional data models (star schema) supporting analytics and reporting use cases
• Implement data quality checks, validation, and monitoring across both batch and streaming pipelines
• Troubleshoot and resolve production issues across Glue jobs, streaming jobs, and Power BI datasets, including root-cause analysis
• Document data flows, pipeline architecture, and data models to support internal knowledge sharing
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field
• 5–7 years of experience in data engineering or ETL development
• Strong, hands-on experience with Apache Spark (PySpark) — transformations, joins, and aggregations at scale
• Strong, hands-on experience with AWS Glue: jobs, crawlers, the Data Catalog, and the surrounding S3 and Athena stack
• Hands-on experience with Spark streaming in production
• Strong, hands-on experience with Power BI: data modeling, DAX, and Power Query
• Strong SQL, including query optimization against large datasets
• Solid understanding of dimensional data modeling (star schema) for analytics and reporting use cases
• Strong debugging and production-support skills across data pipelines
• Excellent written and verbal communication skills for cross-functional collaboration
AI Knowledge & AI-Assisted Development (Required)
• Daily, practical use of AI coding/assistant tools (e.g., Claude Code, GitHub Copilot, Cursor, or similar) to accelerate PySpark, SQL, and DAX development
• Able to critically review and validate AI-generated code, queries, and transformations for correctness, performance, and data integrity before deployment
• Understanding of secure and compliant AI tool usage, including never entering PHI, member data, or other sensitive information into prompts or external AI tools
Preferred Qualifications
• Experience in the health insurance or payer domain: claims, eligibility, enrollment, provider, or member data, with HIPAA-aware data handling practices
• Exposure to IBM App Connect Enterprise (ACE), IBM Integration Bus / Message Broker, or IBM MQ as an upstream source of event and file feeds
• Exposure to healthcare data standards (X12 EDI, HL7, FHIR)
• Experience with additional AWS data services: Redshift, Redshift Spectrum, or EMR
• Relevant certifications: AWS Certified Data Engineer, or Microsoft PL-300 (Power BI Data Analyst)
Senior Data Engineer · Tranzeal Inc.