Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Member of Technical Staff (ML Eval)

Nema
πŸ‡ΊπŸ‡Έ United States
On-site
Staff / Principal
8 months ago
$120,000 – $250,000 / year
  • AI

Not enough detail in this posting to match

Role Description

We're building AI agents that understand how safety-critical hardware gets built and we need your expertise.

In this role you'll own the LLM evaluation infrastructure. You'll build evals that catch regressions before they impact customers, and you'll work directly with the product team to define what "good" looks like for Nema Agent-generated requirements, test cases, and documentation.

You will:

  • Build and maintain evaluation pipelines for LLM outputs

  • Design domain-specific benchmarks for hardware engineering workflow

  • Instrument our AI features to collect human feedback and measure real-world performance

  • Run experiments comparing model architectures and context management strategies

  • Work with customers to understand failure modes and build evals that catch them

  • Do whatever it takes to help the company win

You have:

  • A bias to action

  • Experience building evaluation infrastructure

  • The patience to label data yourself when needed

You are NOT:

  • Interested in publishing papers more than shipping product

  • Expecting a clean dataset to magically appear

Bonus points for:

  • Experience evaluating LLMs in regulated industries (aerospace, defense, automotive, medical)

Member of Technical Staff (ML Eval) Β· Nema

Auto apply with Likeremote