
Member of Technical Staff (ML Eval)
- AI
Not enough detail in this posting to match
Role Description
We're building AI agents that understand how safety-critical hardware gets built and we need your expertise.
In this role you'll own the LLM evaluation infrastructure. You'll build evals that catch regressions before they impact customers, and you'll work directly with the product team to define what "good" looks like for Nema Agent-generated requirements, test cases, and documentation.
You will:
Build and maintain evaluation pipelines for LLM outputs
Design domain-specific benchmarks for hardware engineering workflow
Instrument our AI features to collect human feedback and measure real-world performance
Run experiments comparing model architectures and context management strategies
Work with customers to understand failure modes and build evals that catch them
Do whatever it takes to help the company win
You have:
A bias to action
Experience building evaluation infrastructure
The patience to label data yourself when needed
You are NOT:
Interested in publishing papers more than shipping product
Expecting a clean dataset to magically appear
Bonus points for:
Experience evaluating LLMs in regulated industries (aerospace, defense, automotive, medical)
Member of Technical Staff (ML Eval) Β· Nema