Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
Fundamental logo

Evaluations Team Lead

Fundamental
  • 🇮🇱 Israel
  • Remote
  • Staff / Principal
  • 23 hours ago
  • AI
  • Nexus
  • calibration
  • Temporal
  • Python
  • SQL
  • scikit-learn
  • Snowflake
  • Databricks
  • Equity
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Fundamental

Fundamental is an AI research lab pioneering the future of enterprise decision-making. Our flagship model, NEXUS is the world's most powerful Large Tabular Model (LTM) - purpose-built for the structured records that contain trillions of dollars in business value. With $275m in funding from leading investors and trusted by Fortune 100 companies, Fundamental is giving businesses the Power to Predict.

At Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.

About the role

NEXUS is already in production, and we are working to substantially improve its predictive quality, latency, and cost efficiency. You will lead the team that owns how NEXUS is measured, building the shared evaluation platform that engineering, research, and Applied AI all rely on: consistent benchmarks, curated datasets, and standards for metrics, data splits, leakage prevention, and benchmark contamination that hold up under scrutiny.

Research needs to trust results before committing compute to it, engineering needs regressions caught before they reach production, and when a customer's own data science team benchmarks NEXUS against their own models, your platform is what Applied AI relies upon. You will benchmark NEXUS against competing approaches, using fair tuning budgets, data access, and latency measurement protocols, and turn what you find into research priorities and release recommendations.

This is a player-coach role: you will hire and manage a small team while staying hands-on with the code, the experiment design, and the methodology yourself. There is no evaluation function to inherit here - what you build becomes the standard the rest of the company measures NEXUS against.

Key responsibilities

  • Build a shared evaluation platform for engineering, research, and Applied AI. Continuously add and maintain models and curated datasets so teams can run benchmarks and investigate results independently.

  • Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.

  • Build reproducible pipelines with versioned inputs and artifacts, and integrate regression checks into research and release workflows.

  • Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.

  • Support Applied AI’s customer POC evaluations with tooling, methodological guidance, and analysis.

  • Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics.

  • Turn findings into research priorities, release recommendations, and evidence-backed customer improvement plans.

  • Hire and develop the team, set priorities, and stay hands-on with code and experimental design.

Must have

  • Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.

  • Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.

  • Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits.

  • Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.

  • Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement.

  • Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.

  • Clear written and spoken communication with researchers, engineers, and customer data scientists, including the willingness to challenge claims the evidence does not support.

Nice to have

  • Experience evaluating tabular foundation models or AutoML systems.

  • Experience measuring how model optimisations affect predictive quality and inference performance.

  • Experience with relational or multi-table data, and Snowflake or Databricks environments.

  • Experience evaluating automated or agent-driven ML workflows, including failures that aggregate metrics can hide.

Benefits

  • Competitive compensation with salary and equity

  • Comprehensive health coverage for you and your dependents

  • Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys

  • Relocation support for employees moving to join the team in one of our office locations

  • A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

Evaluations Team Lead · Fundamental

Auto apply with Likeremote