Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior Data Science Engineer

EPAM Systems
๐Ÿ‡ง๐Ÿ‡ท Brazil | ๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ฑ Chile | ๐Ÿ‡จ๐Ÿ‡ด Colombia
Remote
Senior
1 day ago
  • Snowflake
  • Databricks
  • Pandas
  • RBAC
  • Python
  • SQL
  • Git
  • Data Modeling
  • scikit-learn
  • RAG
  • GCP
  • AWS
  • Azure
  • AI
  • Claude Code
  • Cursor
  • BigQuery
  • Large Language Models
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aSenior Data Science Engineer to build reusable data-sharing adapters and governed access patterns that connect a cloud lakehouse to external analytics platforms while supporting reliable ML and LLM-enabled use cases. You will design scalable integrations, ensure secure tenant-scoped access, and deliver production-ready pipelines and tooling.

Responsibilities

  • Design a dual-format lakehouse write layer that supports Delta and Iceberg metadata on one physical dataset
  • Build and validate ingestion pipeline patterns from object storage to an analytical warehouse for structured operational data
  • Implement change-data-capture patterns with Kafka for real-time and near-real-time movement into the lakehouse
  • Develop dependency-aware bookkeeping and data lineage tracking patterns across data pipelines
  • Ensure adapter code is modular, version-controlled, and reusable across new data source integrations
  • Configure external table definitions for governed data products in Snowflake catalogs
  • Validate zero-copy access from Snowflake to Iceberg and Delta tables without data movement
  • Implement tenant-scoped access controls compatible with Snowflake governance metadata
  • Implement and certify a Delta Sharing adapter for zero-copy sharing to Databricks consumers
  • Configure Delta Sharing endpoint registration and sharing agreement management
  • Validate Databricks read access via Delta Sharing for pandas, Spark, and other compatible consumers
  • Test end-to-end freshness and sharing latency against agreed SLA targets
  • Register external connector types and implement auditable RBAC and tenant-scoped authorization for all data-out paths
  • Implement metering hooks compatible with the billing framework for governed data-out flows

Requirements

  • 3+ years of data science and ML engineering experience using Python and SQL
  • Strong leadership skills to drive technical decisions and delivery across integrations
  • Proven project experience delivering production data pipelines and reusable adapters
  • Advanced Python skills with clean, testable code and Git-based workflows
  • Strong data engineering skills with pandas, data modeling, and warehouse/lakehouse concepts
  • Hands-on ML skills with scikit-learn, experiment tracking, and model monitoring (drift, performance)
  • Solid LLM and RAG skills including prompt engineering, embeddings, chunking strategies, and evaluation methods
  • Strong software engineering skills in modular design, debugging, and basic system/API design with latency tradeoffs
  • Working cloud fundamentals across GCP, AWS, or Azure for data and ML workloads
  • Proficiency with AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
  • Upper-Intermediate English proficiency (B2)
  • Strong communication skills to document integration patterns and align with stakeholders

Nice to have

  • Google Cloud Platform experience with BigQuery, GCS, and lakehouse patterns
  • Large Language Models (LLM) experience with API integration, rate limits, and cost management
  • Vector database experience for RAG implementations, including indexing and retrieval tuning

Senior Data Science Engineer ยท EPAM Systems

Auto apply with Likeremote