Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead MLOps Engineer

EPAM Systems
๐Ÿ‡ง๐Ÿ‡ท Brazil | ๐Ÿ‡ฆ๐Ÿ‡ท Argentina
Remote
Staff / Principal
20 hours ago
  • MLOps
  • calibration
  • AI
  • Temporal
  • Configuration Management
  • Python
  • SQL
  • Snowflake
  • pgvector
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aLead MLOps Engineer to head an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one โ€” every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.

This role sets the technical direction for the ML and signal components, ensuring they are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.

The Lead MLOps Engineer will define standards, mentor engineers, and act as the primary technical authority for all MLOps practices across both franchise workstreams.

Responsibilities

  • Define the overall MLOps architecture and strategy for the platform, establishing standards for reproducibility, lineage, and model lifecycle management across both franchises
  • Own Signal Catalogue operations end-to-end: architect signal refresh orchestration triggered by feed read-position advances, define grain translation policies between per-test, per-area, and per-run signal families, and establish provenance capture standards across all 8 signals
  • Set the strategy for Semantic Vector Index versioning: partner with the Lead AI Developer on model and dimension stamp conventions; design and govern the controlled reindex path when the enterprise AI gateway model changes
  • Architect, deliver, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runner
  • Establish and enforce holdout patch-set discipline, define configuration sweep methodology over route limits, thresholds, weights, and Composition setting; oversee catch-rate vs. scope analysis per candidate configuration and approve winning configurations as suggested-weight proposals into the Settings View
  • Lead Model Generation & Experimentation strategy: define the systematic experimentation framework over scoring strategy configurations, guiding the team on which signal weights and route combinations yield the best catch-rate vs. scope trade-off
  • Govern model lineage across configuration versions for both franchises, ensuring cross-franchise consistency and auditability
  • Design and oversee the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channels
  • Define exploration cadence policy: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendations
  • Author and own operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex procedures
  • Mentor engineers on MLOps best practices, review technical designs, and represent the MLOps function in cross-team architecture discussions with data engineering, AI development, and product stakeholders

Requirements

  • 5+ years of experience in MLOps or ML platform engineering, with a proven track record of leading technical initiatives end-to-end
  • Deep expertise in ML model lifecycle management, including versioning, configuration management, and rollback, with experience defining organization-wide standards
  • Strong background in architecting signal computation pipelines, covering scheduled refresh, provenance capture, and grain translation
  • Advanced proficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurement
  • Expert knowledge of as-of temporal data systems or back-test harness design and operation
  • Advanced skills in Python and SQL for ML pipeline automation, with experience setting coding and design standards for a team
  • Deep competency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift counters
  • Proven ability to lead cross-functional collaboration with data engineers and AI developers on feature alignment and shared roadmaps
  • Strong track record of authoring calibration procedures, signal definitions, and operational runbooks that scale across teams
  • Solid experience with Spec Driven Development, ideally as a methodology advocate or champion
  • Demonstrated mentorship and technical leadership experience, including code review, design review, and guiding mid-to-senior engineers
  • Excellent written and verbal communication skills in English (B2+ level)

Nice to have

  • Familiarity with Snowflake ML or Snowpark in production settings
  • Proven showcase of embedding model versioning and controlled reindex orchestration at scale
  • Operational leadership with pgvector or vector store management
  • Background in gaming domain or QA tooling

Lead MLOps Engineer ยท EPAM Systems

Auto apply with Likeremote