DL
Data Engineer Sage Maker
Diverse Lynx India
๐ฎ๐ณ India
On-site
Mid level
2 months ago
- AWS SageMaker
- AWS Cloud
- PySpark
- AWS Glue
- Apache Airflow
- AI/ML
- AWS
- ETL
- ELT
- Athena
- Redshift
- EMR
- IAM
- Python
- Data Modeling
- SQL
- Git
- CI/CD
2 months ago
Job Summary:
We are looking for an experiencedData Engineer with strong expertise inAWS cloud technologies, PySpark, AWS Glue, Apache Airflow, and Amazon SageMaker to design, develop, and maintain scalable, secure, and high-performance data platforms. The ideal candidate will have experience building end-to-end data pipelines, implementing Medallion Architecture, and supporting analytics and AI/ML workloads in a cloud-native environment.
Key Responsibilities
- Design, develop, and maintain scalable, reliable, and fault-tolerant data pipelines on AWS.
- Build and optimize ETL/ELT workflows usingPySpark,AWS Glue, and other AWS services.
- ImplementMedallion Architecture (Bronze, Silver, Gold) for structured and semi-structured data processing.
- Develop and orchestrate data workflows usingApache Airflow, including DAG creation, scheduling, dependency management, and monitoring.
- Build data ingestion and transformation pipelines from multiple data sources into enterprise data lakes and data warehouses.
- Work extensively withAWS services includingS3, Glue, Athena, Redshift, Lambda, EMR, and related cloud-native technologies.
- Enable AI/ML use cases by building feature engineering pipelines and preparing datasets usingAmazon SageMaker.
- Ensure data quality, validation, monitoring, lineage, and observability across the data platform.
- Optimize data storage, partitioning, file formats, and processing performance for cost-efficient and scalable solutions.
- Implement data security, governance, IAM policies, encryption, and compliance best practices.
- Collaborate with Data Scientists, Data Analysts, Application Developers, and Business teams to deliver high-quality data solutions.
- Participate in architecture discussions, code reviews, and performance optimization initiatives.
- Troubleshoot production issues, monitor pipeline health, and ensure high availability of data platforms.
- Strong experience withPython andPySpark.
- Hands-on experience withAWS Glue,Amazon SageMaker,S3,Athena,Redshift,Lambda, andEMR.
- Experience implementingMedallion Architecture and modern Data Lake/Data Warehouse solutions.
- Strong knowledge ofApache Airflow and DAG development.
- Experience with ETL/ELT design, data modeling, and data integration.
- Good understanding of SQL, data partitioning, schema evolution, and performance tuning.
- Experience with version control tools such asGit and CI/CD practices.
- Knowledge of data governance, security, IAM, encryption, and cloud best practices.
Data Engineer Sage Maker ยท Diverse Lynx India