Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

MLOps Engineer

TinyFish
🇺🇸 United States
Hybrid
16 months ago
  • AI
  • Devops
  • Terraform
  • Delta Lake
  • Apache Iceberg
  • Great Expectations
  • MLflow
  • Apache Airflow
  • CI/CD
  • IaC
  • Prometheus
  • Grafana
  • Datadog
  • IAM
  • SOC2
  • GDPR
  • HIPAA
  • AWS
  • GCP
  • Azure
  • Python
  • Docker
  • Kubernetes
  • Kubeflow
  • Weights & Biases
  • Git
  • Ray
  • Flink
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Position Overview

As the first dedicatedML Ops Engineer, you’ll own the tooling and infrastructure that make our ml engineers wildly productive and ensure we are able to efficiently iterate on ML models, prompts, and datasets and deploy our AI systems into a predictable production environment. You’ll bridge the gap between research and DevOps—designing reproducible dataset pipelines, automated experiment workflows, and Terraform-based cloud deployments that scale.


Key Responsibilities

Dataset Management

• Design version-controlled data pipelines (feature stores, data registries) using tools such as Delta Lake, Apache Iceberg
• Implement systems for data validation, lineage tracking, and automated quality checks (e.g., Great Expectations).

Experiment Execution & Tracking

• Build and maintain experiment orchestration with platforms like MLflow, torchx, and Apache Airflow.
• Provide templated systems and tools to ML Engineers that easily launch training/evaluation data processing systems
• Automate hyper-parameter sweeps and A/B tests, exposing clear dashboards for results.

CI/CD

Models/Agents

• workflows that package, test, and promote models and agents through staging to production.
• Implement canary deployments and rollbacks for models/agents services

Terraform Infrastructure-as-Code•

• Author and maintain Terraform modules for all ML infra—networking, GPU/TPU clusters, object storage, secrets, monitoring.
• Enforce best practices for state management, workspaces, and automated plan/apply stages via CI.

Observability & Reliability

• Integrate logging, tracing, and metric collection (Prometheus, Grafana, Datadog) across data pipelines and model endpoints.
• Set SLIs/SLOs for data freshness and model latency; implement alerts and runbooks.

Security & Compliance• Work with Security to implement IAM least-privilege, key rotation, and data-encryption policies.
• Support audit requirements (SOC 2, GDPR, HIPAA where applicable).


Minimum Qualifications

  • 5+ years combined experience in DevOps, Data Engineering, or ML Ops roles.

  • Strong Terraform skills; ability to craft reusable modules and navigate complex state.

  • Production experience with at least one cloud provider (AWS, GCP, or Azure).

  • Proficiency inPython and containerization (Docker); familiarity with Kubernetes or serverless batch systems.

  • Hands-on knowledge ofML experiment platforms (MLflow, Kubeflow, Weights & Biases, or similar).

  • Experience with workflow execution frameworks (Kubeflow, Apache Airflow)

  • Understanding of moderndata-versioning/feature-store concepts and tools.

  • Solid grasp ofCI/CD principles, Git workflows, and infrastructure testing.

  • Excellent communication skills—capable of partnering with Data Scientists, Software Engineers, and Security teams.

Preferred (Nice-to-Have)

  • Experience withGPU orchestration (NVIDIA DGX, Karpenter, or Ray).

  • Familiarity withIaC security scanning (Checkov, tfsec).

  • Exposure topolicy-as-code (OPA/Gatekeeper).

  • Prior work inreal-time streaming (Kafka, Flink) andonline feature serving.

  • Contributions to open-source ML Ops projects.


Reporting Structure

Reports to: Director of Infra

MLOps Engineer · TinyFish

Auto apply with Likeremote