Senior Serverless Spark Migration Engineer
- AI
- Apache Spark
- Hadoop
- AWS
- GCP
- EMR
- BigQuery
- PySpark
- Scala
- SQL
- Yarn
- Hive
- CI/CD
- IaC
- Terraform
- Devops
- FinOps
- Git
- Delta Lake
- Airflow
- Docker
- Kubernetes
Senior Platform Engineer, Cloud Infrastructure
Type: Remote (Brazil, Mexico)
Coverage: Pacific Hours (8:00 AM – 5:00 PM PST)
About Virtasant
Virtasant is a global cloud and technology services company helping organizations modernize, optimize, and build at scale. We work with enterprise customers on complex cloud, data, infrastructure, and AI initiatives, bringing together deep technical expertise and hands-on delivery.
About the Role
We’re looking for aSenior Serverless Spark Migration Engineer to help modernize a large-scale enterprise data platform.
You’ll lead the migration of productionApache Spark workloads from on-premise Hadoop/Spark environments to cloud-native and serverless architectures across AWS and GCP. This is a hands-on engineering role spanning workload assessment, architecture, application refactoring, migration execution, performance optimization, automation, and production readiness.
The goal is not simply to lift and shift existing workloads. You’ll determine the right target architecture for each workload and establish repeatable patterns that can eventually support migration at significant enterprise scale.
What You’ll Do
Lead migrations of enterprise Spark workloads from on-premise environments toAWS and GCP.
Assess Spark applications, clusters, configurations, dependencies, data flows, and resource utilization.
Determine the right migration approach acrossrehost, replatform, refactor, modernize, or retire.
Modernize traditional cluster-based workloads forserverless Spark where appropriate.
Design and implement architectures using technologies such asAWS EMR Serverless, S3, Glue, Lake Formation, GCP Dataproc Serverless, GCS, and BigQuery.
Refactor legacyPySpark/Scala/Spark SQL applications for cloud portability, scalability, and reliability.
Migrate workloads from environments usingHadoop, HDFS, YARN, Hive, and on-prem Spark clusters.
Troubleshoot and optimize Spark workloads across partitioning, shuffle behavior, joins, data skew, execution plans, executor configuration, serialization, and SQL execution.
Benchmark performance and optimize serverless workloads forperformance, reliability, and cloud cost.
Build reusable migration tooling, automation, templates, and frameworks.
Implement CI/CD and Infrastructure as Code using tools such asTerraform.
Define testing, validation, cutover, rollback, observability, and production-readiness patterns.
Partner with Data Engineering, ML, Cloud Architecture, Platform Engineering, DevOps/SRE, Security, Governance, and FinOps teams.
What We’re Looking For
8+ years of experience across data engineering, distributed systems, cloud engineering, or platform engineering.
5+ years of hands-on Apache Spark experience in enterprise environments.
StrongPySpark and/or Scala development experience.
Proven experience migratinglarge-scale Spark workloads between infrastructure platforms.
Hands-on experience withboth AWS and GCP.
Experience withon-premise Hadoop/Spark ecosystems, including technologies such as HDFS, YARN, and Hive.
Deep understanding ofSpark internals and distributed processing.
Strong SQL and data engineering fundamentals.
Experience with cloud data lakes and object storage.
Strong production troubleshooting and performance-tuning experience.
Experience withCI/CD, Git, and Infrastructure as Code.
Ability to own migration work end-to-end, from discovery and architecture through production cutover and optimization.
Nice to Have
Experience withEMR/EMR Serverless, Dataproc/Dataproc Serverless, Glue, Lake Formation, BigQuery, Delta Lake, Iceberg, Kafka, Airflow, Terraform, Docker, or Kubernetes is valuable.
What Success Looks Like
You can take ownership of the complete migration lifecycle:
Discover → Assess → Design → Refactor → Migrate → Validate → Optimize → Operate
You understand both the legacy Hadoop/Spark world and modern cloud-native data platforms, and can make pragmatic architecture decisions based on workload characteristics rather than simply reproducing an existing environment in the cloud.
Senior Serverless Spark Migration Engineer · Virtasant