Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead Data Science Engineer

EPAM Systems
๐Ÿ‡ง๐Ÿ‡ท Brazil | ๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡จ๐Ÿ‡ฑ Chile | ๐Ÿ‡จ๐Ÿ‡ด Colombia
Remote
Staff / Principal
2 days ago
  • AI
  • Snowflake
  • Databricks
  • RBAC
  • Python
  • Pandas
  • RAG
  • SQL
  • Git
  • GCP
  • AWS
  • Azure
  • Windows
  • Claude Code
  • Cursor
  • BigQuery
  • Large Language Models
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aLead Data Science Engineer to design and deliver reusable data sharing adapters and governed access patterns across a cloud lakehouse and external data platforms. You will guide high-impact architecture decisions, ensure reliable data and AI workflows, and help the team ship secure integrations.

Responsibilities

  • Design a reusable lakehouse write layer with dual-format metadata to support multiple consumers
  • Build and validate ingestion pipeline patterns from object storage to an analytical warehouse for structured operational data
  • Implement change-data-capture patterns for real-time and near-real-time data movement into the lakehouse
  • Develop dependency-aware bookkeeping and data lineage tracking patterns across pipelines
  • Ensure adapter code is modular, version-controlled, tested, and reusable across new integrations
  • Configure external table definitions for shared lakehouse data products in Snowflake catalogs
  • Validate zero-copy read access from Snowflake to shared Iceberg and Delta tables without data movement
  • Implement tenant-scoped access controls aligned with external catalog governance requirements
  • Implement and certify a Delta Sharing adapter for live data sharing to Databricks consumers
  • Configure Delta Sharing endpoints and manage sharing agreements for multiple tenants
  • Validate consumer access via supported clients while meeting freshness and latency expectations
  • Register connector types in a governed connector registry and enforce auditable RBAC on data-out paths
  • Implement metering hooks compatible with governed billing requirements for external data flows

Requirements

  • 5+ years of data science or ML engineering experience with production Python and pandas
  • Experience building RAG applications using embeddings and retrieval pipelines
  • Experience writing SQL for analytical data workflows and validation
  • Strong technical leadership skills to drive architecture decisions and mentor peers
  • Proven project delivery skills across multi-system data integration workstreams
  • Solid software engineering skills in modular design, testing, debugging, and Git workflows
  • Hands-on cloud platform skills with GCP, AWS, or Azure fundamentals
  • Strong LLM fundamentals knowledge including tokenization, attention, context windows, and sampling
  • Practical prompt engineering skills with structured outputs and few-shot techniques
  • Robust evaluation and monitoring skills including drift, performance, and hallucination detection
  • Strong communication and collaboration skills across engineering and data stakeholders
  • Upper-Intermediate English proficiency (B2, Upper-Intermediate)
  • Active experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor

Nice to have

  • Google Cloud Platform experience with BigQuery and object storage patterns
  • Large Language Models (LLM) API integration experience including streaming, rate limits, and cost controls
  • Vector database experience with indexing, chunking strategies, and retrieval tuning
  • Experience optimizing LLM latency and cost using caching, batching, and model routing

Lead Data Science Engineer ยท EPAM Systems

Auto apply with Likeremote