Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
Hub logo

Machine Learning Engineer

Hub
  • 🇫🇷 France
  • On-site
  • 2 hours ago
  • $90,000 – $120,000
  • AI
  • PostgreSQL
  • Qdrant
  • Core ML
  • calibration
  • Language Models
  • AWS
  • Computer Vision
  • AWS S3
  • Kubernetes
  • Temporal
  • GitHub
  • Hugging Face
  • ROS
  • PyTorch
  • CUDA
  • vLLM
  • Equity
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Hub

Real-world training data for frontier AI labs and robotics.

Tech

## Our Tech Stack & Engineering Challenges At Hub, we build and operate distributed, multimodal infrastructure that orchestrates large-scale real-world data pipelines from collection to enterprise delivery. ### The Core Pipeline Our engineering team builds and scales infrastructure across four core layers: * **Intelligent Distribution (HubApp):** Algorithmic matching engines that dispatch collection tasks to a verified network of global contributors based on hardware, location, and language profiles. * **Automated ML Processing:** Modality-specific pipelines for audio and egocentric video, including voice activity detection, denoising, speaker diarization, synchronization, motion analysis, and contextual labeling. * **HITL QA Engine:** Human-in-the-loop workflows that route processed assets to human verification nodes to ensure scope-of-work compliance. * **Data & Vector Architecture:** Petabyte-scale data ingestion using object storage, PostgreSQL for metadata, and Qdrant for semantic indexing and asset retrieval. ### Hard Problems We Are Solving * **The Robotics Frontier:** Building data pipelines for embodied AI using real-world first-person environmental and motion data to train advanced robot policies. * **High-Throughput Refactoring:** Converting experimental data science workflows into scalable, containerized production systems. * **Absolute Traceability:** Maintaining end-to-end data provenance and integrity from asset capture through delivery to AI models.

The role

Turn raw multi-camera footage into the 3D, verified datasets that frontier robotics labs train on.### About HubHub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.### The roleYou own Hub's ML pipeline end to end, from raw multi-camera capture in the field to the dataset a frontier robotics lab trains on, side by side with our core ML team. You are the technical counterpart for each lab program, and your pipeline decides what we are allowed to ship.### What you'll own\- 3D and stereo vision on capture rigs we build ourselves: multi-camera RGB, RGB-D and IMU, calibration, stereo and metric depth, SLAM and trajectories in a world frame, 3D reconstruction of the scene and the manipulation.\- Hand tracking across synchronised cameras and fine-grained manipulation, plus annotation at scale against demanding customer taxonomies. Every human verdict becomes a training label.\- Quality control as an ML problem: vision-language models as judges, thresholds per customer, human reviewers where models fall short. You own the eval sets and the call on what we trust.\- Customer pipelines on AWS: extend our shared modules, build what a new spec demands, and keep throughput and cost per processed hour under control.\- Egocentric video with narration across languages: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.\- The technical relationship with each lab program: specs, delivery format, first samples and feedback loops.\- What comes next: tactile and teleoperation data today, and robot learning on our own data in the mid term.### You might be a fit if\- 3+ years of applied ML in production, in computer vision, 3D vision or robotics perception, ideally from a top engineering school. Less experience is fine for outliers: the bar is what you've built.\- 3D and stereo vision in practice: multi-view geometry, intrinsics and extrinsics, stereo depth, SLAM or visual-inertial odometry, 3D reconstruction.\- CV models you trained and deployed in production: detection and tracking, pose and hand estimation, depth, action recognition.\- ML infrastructure on AWS: S3, GPU instances, Kubernetes, distributed training and inference, with throughput and cost you measure.\- VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.\- The video stack: codecs, frame timing and temporal alignment across sensors.\- Robotics knowledge: you understand how robot learning uses this data, from imitation learning to VLAs and teleoperation.\- An active GitHub or Hugging Face, and you follow the literature well enough to tell what's worth implementing from what's noise.\- Agentic engineering as a craft: a custom harness, and agents that verify their own work through tests, training runs and evals.### Nice to have\- Egocentric vision, IMUs, MCAP, ROS.\- World models, video generation, VLAs or robot foundation models.\- An applied PhD or published work in 3D vision or robot learning.### Stack\- PyTorch, CUDA, multi-GPU training and distributed inference. Throughput matters as much as accuracy.\- Multi-camera geometry: intrinsics, camera-to-IMU extrinsics, fisheye rectification, stereo depth, SLAM and trajectories in a world frame.\- Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.\- Frontier VLMs and speech models behind one interface, plus open-weight judges we serve with vLLM and fine-tune on our own data.\- AWS: S3 for bytes, GPU instances and Kubernetes for compute, Postgres for state. Cost per processed hour is an engineering target.\- Every threshold traces to a customer requirement. Every quarantine carries a code, evidence and an owner.### Why Hub\- Small core team by design, San Francisco pace. Urgent things get handled when they come up.\- One owner per project, with 1 to 3 numbers that say whether it's working.\- Everyone runs AI agents, not only engineers. Uncapped frontier models, shared skills and knowledge base.\- Constant, proactive communication: Slack, huddles, short stand-ups. No black box.\- Team across San Francisco, São Paulo and Paris, where our research lab is opening.### What we offer\- $90,000 to $120,000 yearly salary.\- Stock options ranging 0.05 - 0.2%.\- Your own GPU budget.\- Based in Paris, in our office opening soon. Hybrid: ideally most days on site, at least one day a week (or one week a month if you live outside Paris).\- Direct work with the founders and with the biggest AI labs.### How we hire1\. A short application read by the team. We open your GitHub and Hugging Face first.2\. A technical conversation with one of our ML engineers.3\. A build task on real data from our pipeline. We watch how you break the problem down, what you measure and how you verify your own work.4\. A final conversation with a founder and an ML engineer.5\. An answer within 48 hours.### How to applySend a short application in your own words:\- Your CV\- The achievement you're most proud of\- Your GitHub and Hugging Face links\- The story behind it

Machine Learning Engineer · Hub

Auto apply with Likeremote