Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior AI Engineer with Spark, AWS Services

EPAM Systems
πŸ‡ΊπŸ‡Έ United States | πŸ‡¦πŸ‡² Armenia | πŸ‡°πŸ‡Ώ Kazakhstan | πŸ‡°πŸ‡¬ Kyrgyzstan | πŸ‡ΊπŸ‡Ώ Uzbekistan
Remote
Senior
1 day ago
  • AI
  • AWS
  • Machine Learning
  • RAG
  • OpenSearch
  • ETL
  • ELT
  • AI/ML
  • Athena
  • Step Functions
  • DynamoDB
  • Python
  • SQL
  • Bedrock
  • API Gateway
  • CloudWatch
  • Docker
  • Snowflake
  • Pinecone
  • CI/CD
  • Jenkins
  • Git
  • Bitbucket
  • Terraform
  • CDISC
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aSenior AI Engineer with Spark and AWS Services expertise to join the RBQM Production Pod within the program. You will build and maintain data pipelines that power AI/GenAI applications for Risk-Based Quality Management in clinical trials. This role focuses on RAG document ingestion, vector indexing, and building data APIs for AI applications.

Responsibilities

  • Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data
  • Build and manage vector databases (AWS OpenSearch) for RAG-powered AI workflows
  • Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports)
  • Build and expose data APIs for AI application consumption
  • Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures
  • Manage data quality, lineage, and governance for AI/ML data pipelines
  • Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)
  • Collaborate with Data Scientists and Backend Developers in an integrated pod team

Requirements

  • 5+ years of hands-on data engineering experience at scale
  • Expertise in RAG document ingestion pipelines (chunking, embedding, vector indexing)
  • Proficiency in AWS OpenSearch as a vector database for RAG workflows
  • Advanced proficiency in Python, including SQL and Spark SQL
  • Skills in unstructured data transformation (PDF, DOCX) for RAG/LLM applications
  • Familiarity with AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB
  • Knowledge of containerization with Docker
  • Capability to build custom pipelines from scratch, beyond configuring out-of-the-box services
  • Proficiency in English at a B2+ level

Nice to have

  • Background in pharmaceutical or life sciences domain
  • Familiarity with Snowflake, Pinecone (vector DB alternative)
  • Knowledge of SageMaker processing jobs
  • Skills in CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure tools (CDK or Terraform)
  • Understanding of clinical data standards (CDISC, ADaM, SDTM)

Senior AI Engineer with Spark, AWS Services Β· EPAM Systems

Auto apply with Likeremote