Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior MLOps Engineer

EPAM Systems
๐Ÿ‡ฆ๐Ÿ‡ท Argentina | ๐Ÿ‡ง๐Ÿ‡ท Brazil
Remote
Senior
22 hours ago
  • MLOps
  • calibration
  • AI
  • Temporal
  • Configuration Management
  • Python
  • SQL
  • MLflow
  • Kubeflow
  • Snowflake
  • pgvector
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aSenior MLOps Engineer to join an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one โ€” every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.

This role ensures the ML and signal components are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.

Responsibilities

  • Own Signal Catalogue operations: signal refresh orchestration triggered by feed read-position advances, grain translation between per-test, per-area, and per-run signal families, and provenance capture across all 8 signals
  • Operate the Semantic Vector Index versioning: coordinate with the Senior AI Developer on model and dimension stamp conventions; design and execute the controlled reindex path when the enterprise AI gateway model changes
  • Design, implement, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runner
  • Enforce holdout patch-set discipline, configuration sweep over route limits, thresholds, weights, and Composition setting; produce catch-rate vs. scope tables per candidate configuration and publish winning configurations as suggested-weight proposals into the Settings View
  • Lead Model Generation & Experimentation: systematic experimentation framework over scoring strategy configurations, tracking which signal weights and route combinations yield the best catch-rate vs. scope trade-off
  • Maintain model lineage across configuration versions for both franchises
  • Implement the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channels
  • Manage exploration cadence support: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendations
  • Contribute to operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex procedures

Requirements

  • 3+ years of experience in MLOps or ML platform engineering
  • Expertise in ML model lifecycle management, including versioning, configuration management, and rollback
  • Background in signal computation pipeline design, covering scheduled refresh, provenance capture, and grain translation
  • Proficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurement
  • Knowledge of as-of temporal data systems or back-test harness design and operation
  • Skills in Python and SQL for ML pipeline automation
  • Competency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift counters
  • Capability to collaborate with data engineers and AI developers on feature alignment
  • Qualifications in documenting calibration procedures, signal definitions, and operational runbooks
  • Familiarity with Spec Driven Development
  • English proficiency at an Upper-Intermediate level (B2) or higher

Nice to have

  • Understanding of MLflow, Kubeflow, or equivalent experiment tracking platforms
  • Familiarity with Snowflake ML or Snowpark
  • Showcase of embedding model versioning and controlled reindex orchestration
  • Skills in pgvector or vector store operational management
  • Background in gaming domain or QA tooling

Senior MLOps Engineer ยท EPAM Systems

Auto apply with Likeremote