
Senior Manager, AI Foundation Model
Merlin Labs
🇺🇸 United States
Hybrid
Manager or above
2 weeks ago
- AI
- PyTorch
- C++
- IEC
- Unlimited time off
- Pension
2 weeks ago
About You: You have built learned decision-making systems that left the lab and ran on real hardware with real consequences. You are fluent in modern model architecture and post-training, but you are not a benchmark chaser — you have been in the room when a learned system was asked to justify itself to people who sign off on safety, and you know the difference between a model that performs well and a model whose behavior you can characterize. You want to work on a problem where “it works most of the time” is not a result.
Responsibilities:
- Technical strategy: own Merlin's foundation and world-model work — architecture selection, build-vs-adapt decisions, post-training approach, and the capability roadmap that supports it.
- Team leadership: lead and mentor a small team of world-model and post-training engineers; set the technical bar and the review culture for model work across AI Core.
- System interface: design the model interface to the rest of the autonomy stack — structured, schema-constrained plan outputs that a deterministic verifier can accept or reject, never free-form actuator authority.
- Evaluation: define what “good” means before training begins — build the evaluation harness, capability taxonomy, and regression suite that gate every model release, in partnership with the Data/Sim/Release pillar.
- Safety-relevant outputs: establish uncertainty quantification and out-of-distribution detection as first-class model outputs, not afterthoughts — downstream safety monitoring depends on them.
- Benchmarking: deliver an honest, reproducible comparison between learned planning and Merlin's current rule-based behavior planning across representative mission profiles, including the cases where the learned approach loses.
- Certification partnership: work with Systems Engineering, Certification, and the Chief Architect to keep model design inside what is defensible to a regulator, and to shape what “defensible” will mean for learned components.
- Research judgment: track the external research frontier and make disciplined calls about what Merlin adopts, builds, or ignores.
Qualifications:
- Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or a related subject.
- 8+ years building AI systems, with 3+ years leading technical teams or owning a major model program.
- Proven team management experience shipping high-tech, AI-powered models into production — hiring and developing AI engineers, setting technical direction and priorities, and owning delivery from research through deployment.
- Demonstrated ownership of a learned system that shipped into a physical, real-time product — robotics, autonomous vehicles, aerospace, or industrial autonomy.
- Depth in at least two of: world models and learned dynamics; sequence models applied to planning or control; post-training (SFT, preference optimization, RL fine-tuning); structured or constrained generation.
- Rigorous evaluation practice: you have built eval harnesses that caught regressions before customers did, and you can explain why a model's aggregate metric improved while a specific behavior got worse.
- Strong PyTorch; comfortable reading and reasoning about the C++ real-time systems your models feed.
- You write clearly. Architecture decisions here get read by systems engineers, safety engineers, and regulators — not only by other AI engineers.
Nice to Have:
- Experience with learned components in a certified or regulated product (DO-178C, ISO 26262, IEC 62304).
- Background in classical planning, behavior trees, MCTS, or hierarchical task networks — you'll be replacing and interoperating with exactly these.
- Familiarity with aviation domain structure: flight phases, ARINC 424 procedures, ATC phraseology.
- Publications or open-source contributions in embodied AI, world models, or robot learning.
Senior Manager, AI Foundation Model · Merlin Labs