Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Compute Infrastructure Lead

UMA
๐Ÿ‡ซ๐Ÿ‡ท France
On-site
Staff / Principal
2 weeks ago
  • Fabric
  • Elastic
  • Node.js
  • Ray
  • Prometheus
  • Grafana
  • MLflow
  • AI
  • PyTorch
  • Linux
  • InfiniBand
  • Slurm
  • Kubernetes
  • Python
  • Technical Writing
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Your Mission

AsCompute Infrastructure Lead, you will own and scale thecompute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads โ€” reliably, efficiently, and at scale โ€” so our models can go from research to production without the cluster becoming the bottleneck.

This is a hands-on, high-impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero โ€” but you'll have the mandate to shape the architecture that takes usfrom a research cluster to a production-scale, multi-provider fleet, and to production-grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end โ€” multi-provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use โ€” and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.

The technical problem is unusually rich for this stage. We are de-risking a stack built onpre-training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real-time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need โ€” a true multi-provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves โ€” does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role;this role is compute, including how processing jobs actually run on it.

Key responsibilities :

  • Own ourcompute platform end to end โ€” from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first-class goals

  • Build amulti-provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs

  • Design and operate thecloud scheduler โ€” quotas, priority, preemption, topology-aware placement, anddynamic checkpointing so jobs survive node failure, preemption, and provider switches

  • Stand up adistributed compute framework fortraining and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real-time / online-learning path

  • Orchestratedata-processing workloads at scale โ€” CPU and cheaper GPUs, batch and streaming โ€” so post-processing, dataset jobs, and training share one reliable compute fabric instead of ad-hoc scripts

  • Delivervirtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes

  • Buildobservability that works from any provider โ€” system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) โ€” so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days

  • Owncapacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role โ€” but procurement is part of making the infra succeed

  • Help setproduction-grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want

What You Bring to the Table

  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level

  • Proven track recordbuilding and operating infrastructure for large-scale AI model training โ€” not inference-only. Multi-node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long-running jobs healthy at scale

  • Deep, hands-on experience withGPU clouds and cluster operations: provisioning, Linux, high-performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes

  • Built or ownedschedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) โ€” including checkpointing, elasticity, and preemption โ€” so GPUs stay busy and jobs come back from failure

  • Treatobservability, reliability, and cost as core engineering concerns, not afterthoughts

  • StrongPython and systems engineering, with the taste to build tooling that researchers actually want to use

  • Experience working withGPU providers on capacity and commercial terms โ€” you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am

  • Ability toreason about systems end-to-end โ€” performance, scalability, reliability, cost โ€” and make and defend the right trade-offs

  • Thrive in ahands-on, fast-paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI and systems

  • Bonus :online / continuous RL, real-time training loops, or other always-on learning systems

  • Bonus : multi-cloud / multi-provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers,VM-based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes

  • Bonus : robotics, autonomous vehicles, or other embodied/physical-AI training stacks โ€” adjacent large-scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required

  • Bonus : public projects, open-source contributions, maintained tools, or technical writing

  • Wevalue exceptional builders over perfect resumes. If you have a world-class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, we strongly encourage you to apply โ€” even if you don't tick every box. Robotics experience is a plus, not a requirement.

Compute Infrastructure Lead ยท UMA

Auto apply with Likeremote