Subscribe to the latest remote jobs:

Senior Staff Applied ML Engineer - Consumer Agents

Senior Staff Applied ML Engineer - Consumer Agents

from

Senior Staff Applied Machine Learning Engineer, Consumer Agents

Location: Remote — North America

About the role

The Consumer Agent team builds Shopify's consumer-facing agentic shopping experiences — the AI agents inside the Shop app and on merchant storefronts that help millions of buyers search, discover, and buy. As a Senior Staff Applied Machine Learning Engineer, you'll be the person we point at the hardest problems on the surface — distillation to own-GPU inference, agentic memory, whole-page optimization, propensity modeling, or the evaluation rigor that ties them together — and trust to drive it to a shipped, measured outcome. This is a senior individual-contributor seat with deep autonomy: you'll own ambiguous, high-leverage problems end to end, raise the bar for evaluation across the team, and ship weekly on a fast-moving consumer surface. You'll work at the intersection of LLM orchestration, model training and distillation, search and retrieval, and cost/latency engineering, partnering with frontier model labs and ML teams across Shopify.

Key Responsibilities

  • Build and improve student models through supervised fine-tuning and knowledge distillation, including the serving, generation-termination, and throughput-economics work that own-GPU inference demands.

  • Design and raise the bar for ML evaluation across the team — judge design, offline and online metrics, and the judgment to know when a number is real.

  • Develop agentic memory and personalization signals — buyer and user representations, profile generation — that feed predictive layers.

  • Work on whole-page optimization: page composition, ranking, and blending for the agent results surface.

  • Build propensity and predictive models for intent and purchase propensity across agent surfaces.

  • Move fluidly between distillation, memory, ranking, propensity, and evaluation as priorities shift, without lengthy re-tooling.

  • Apply rigorous experimentation on a high-traffic surface — proper holdouts, A/B rigor, and counterfactual or off-policy evaluation where online tests aren't enough.

  • Integrate cleanly with an LLM orchestration stack already in flight — tool calling, structured output, and retrieval-augmented patterns.

  • Become the technical reference point the team routes its hardest model and evaluation calls to.

  • Contribute to cross-team ML architecture conversations across Shopify's ML organization and external model-lab partnerships.

  • Ship weekly, absorbing ambiguity rather than adding coordination overhead.

Qualifications

  • You've shipped LLM systems to production at real scale — systems with real users, not research prototypes.

  • Hands-on depth in model training and distillation: supervised fine-tuning and knowledge distillation in practice.

  • Comfort with the serving and throughput-economics side of inference, not just the training run.

  • Rigorous evaluation methodology: judge design, offline and online metrics, and the judgment to distinguish a real result from an artifact.

  • You operate as a senior individual contributor with high autonomy — you take an ambiguous, hard problem and drive it quickly to a measured outcome without hand-holding.

  • Breadth over a single specialty: you can move across distillation, memory, whole-page optimization, propensity, and evaluation.

  • Strong communication and collaboration in a fast-paced, cross-functional environment.

  • Able to work with significant overlap with North America time zones.

Nice to have

  • E-commerce, marketplace, commerce-search, recsys, or ads background — with instincts for selection bias, position bias, and the gap between offline and online metrics.

  • Inference and serving internals: vLLM or similar engines, structured and constrained decoding, KV-cache and GPU throughput economics, and latency budgets.

  • Reinforcement learning experience — RLHF, RLAIF, or preference optimization.

  • Propensity, ranking, or classical applied ML at scale: gradient boosting, neural ranking, calibration, position-bias correction.

  • Counterfactual evaluation experience — inverse propensity scoring, doubly-robust estimators, off-policy evaluation.

  • Experience working alongside or with frontier AI products and agentic systems.

by @maxrusakovic