
Senior / Staff ML Ops Engineer
Waabi
🇨🇦 Canada | 🇺🇸 United States
Hybrid
Staff / Principal
1 day ago
- AI
- Kubernetes
- Node.js
- Python
- CI/CD
- Helm
- AWS
- IAM
- IaC
- Terraform
- Pulumi
- PyTorch
- LiDAR
- Parquet
- Argo Workflows
- Ray
- Kubeflow
- Slurm
- Bazel
- Health insurance
- Equity
- Unlimited time off
1 day ago
You will..
- Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable.
- Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible.
- Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down.
- Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost.
- Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal sensor data, so jobs saturate GPUs instead of waiting on I/O.
- Turn one-off Python into durable tooling — tested, documented, observable libraries, CLIs, and services with sane defaults, and deletions where they're overdue.
- Make experiments legible, with the teams who live in them: experiment hygiene, dashboards researchers trust, a real model registry, and lineage from dataset to checkpoint to simulation result.
- Ship CI/CD for models alongside autonomy and simulation, so a model change is validated the same way a code change is.
- Build observability across the ML stack — utilization, throughput, failure modes, queue times, cost per experiment. When a job fails at 3am on node 47, the researcher should find out why without you.
- Treat docs, onboarding, and support as product surface — golden-path guides, a new researcher productive on day two, office hours that turn repeat questions into shipped fixes.
- Drive adoption, not just availability. Prototype with real users, watch them work, iterate. A tool nobody adopts didn't ship.
- Make the platform boringly reliable — fewer failures, faster recovery, and none of the manual steps that quietly cost a team days.
- Build guardrails that don't feel like walls, with Security, IT, and Infrastructure: access controls, data handling, and cost governance that hold up in an IP-sensitive environment while staying self-serve.
- 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems.
- Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and the ability to debug a cluster under load rather than restart it.
- Excellent Python, and a track record of designing APIs and CLIs other people enjoy using.
Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar). - Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling — from the perspective of someone who made them pleasant for others to use.
- Fluency with containers, CI/CD, and modern build systems, including large monorepos.
- The ability to influence without authority: evaluate a framework on its merits, pilot it credibly, and persuade skeptical senior engineers to change how they work.
- A collaborative default — you'd rather co-own a system than draw a boundary around your part of it.
- User empathy: you'd rather fix the third-most-interesting problem blocking ten people than the most interesting one blocking nobody.
- Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory.
- Passionate about self-driving technologies and frontier AI, and about what a small, world-class team can do with the right infrastructure.
- Internal developer platform, research platform, or DevEx work — with a story about a tool whose adoption you grew from zero.
- Large-scale distributed GPU training: hundreds to thousands of accelerators, NCCL, high-performance cluster networking, collective communication tuning.
- High-throughput loading of LiDAR or camera data, and formats such as Parquet or WebDataset.
- Workflow and scheduling systems — Argo Workflows, Ray, Flyte, Kubeflow, or Slurm.
- Build-system depth (Bazel or similar), including remote caching in a monorepo.
- Simulation infrastructure or large-scale batch evaluation pipelines.
- Background in ML, robotics, or autonomous systems infrastructure.
- Security- and IP-sensitive production environments.
- Open-source contributions to ML infrastructure or developer tools.
Perks/Benefits:
- Competitive compensation and equity awards.
- Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time employees only).
- Unlimited Vacation.
- Flexible hours and Work from Home support.
- Daily drinks, snacks and catered meals (when in office).
- Regularly scheduled team building activities and social events both on-site, off-site & virtually.
- As we grow, this list continues to evolve!Â
Senior / Staff ML Ops Engineer · Waabi