Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Founding GTM

Openbenchmarks
πŸ‡ΊπŸ‡Έ United States
On-site
1 day ago
$100,000 – $150,000
  • AI
  • Oracle
  • Data Visualization
  • Claude Code
  • Claude
  • SEO
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Openbenchmarks

We help humans & agents pick tools via Independent Benchmarks

The role

**About OpenBenchmarks**Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.Agents increasingly prefer open, independent and grounded benchmarks to make decisions.Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.Our mission is to be the trusted evaluation layer for agents.Founders previously led AI research and Infra teams at Oracle and Appfolio; we started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.**About the role**Benchmarks are how companies market their products to agents. We build those benchmarks.There's a lot of nuance in every benchmark - what gets evaluated, which metrics matter, release cycles, data refreshes. A benchmark isn't static, and a changing benchmark constantly surfaces new insight into how the tools on it actually perform. Turning that into content - posts, graphs, chart is super important to communicate the nuance of the benchmark in the most clear way possible.That's your job.Here are the broad themes that you'll be working on**_Content from benchmarks_** - You'll run benchmarks, query the data, and dig into individual failures to find what's actually there - where a vendor is strong and why, what the failure mode of the field is, what the trade-off costs. You use that to create content: the finding, graphs and more, and what we're changing in v2. Doing the digging deeply yourself is what makes the writing credible.**_Data visualization and Graphs_** - the graphs & charts are how the trade-offs are shown. Precision against latency, F1 against cost etc. You'll design and build these yourself using Claude Code, Claude Design.**_Distribution to people and agents_** - We have two audiences. People on LinkedIn, X, and HN. And the agents that read, index, and cite it when someone asks which tool to pick. You'll own how the work gets found by both - which means the writing has to be genuinely , because slop doesn't survive either audience.**_Knowing & being updated on the benchmark landscape_** - you know what benchmarks already exist - SWE-bench, Terminal-bench, BrowseComp, continuously understand the leaderboards vendors cite in their launch posts - and be able to say precisely where they're saturated, where they're gamed, and where the gap is that we should fill. You'll bring proposals for what we benchmark next, and the case for why it matters.**Preferred** * Background in Statistics, Engineering or Data Science* Working Knowledge of Search Engine Optimization* Have taste and opinions on great design vs bad design.* Working Knowledge of Claude Code, Claude Design or creating data visualizations that stand out.

Founding GTM Β· Openbenchmarks

Auto apply with Likeremote