06-reference

every benchmarks dont know your job

2026-08-25·reference·source: Every (Context Window)·by Katie Parrott
evalsbenchmarksai-adoptionagent-reliabilityharness-engineering

"Benchmarks Don't Know Your Job" — Katie Parrott (Every, Context Window, 2026-08-25)

Why this is in the vault

Sharpest public articulation this month of the "leaderboard score ≠ your task" problem, illustrated with Every's own internal eval (KateBench) rather than a vendor pitch.

The core argument

Companies that spend heavily on AI models often know their spend and public benchmark rankings but not whether a model actually saves time or produces trustworthy work on their specific tasks. Mercor CEO Brendan Foody and Box CEO Aaron Levie ("Enterprises will not be able to go just on vibes") are cited making the same point: public benchmarks show relative capability, not whether a model caught the clause your lawyers care about or preserved your house style. With model choice multiplying (open-source token share on Vercel's platform going 28%→62% in two months; legal-worker Codex adoption up 108x since February per a16z), picking the top-scoring model is no longer a real strategy.

Every's answer is task-specific "offline evals" built from real work. Its flagship internal instrument, KateBench, is an AI copyeditor trained on ~30,000 of editor-in-chief Kate Lee's past edits; it suggests changes in a Google Doc and tracks which ones editors accept, reject, or still rewrite afterward. The piece is candid about the instrument's own failure modes: KateBench's engineer, Jannik Jung, found the tool silently dropped suggestions past a 40-item cap (raised to 80) — meaning a headline 85–90% acceptance rate was partly an artifact of undercounting, not quality — and that acceptance rates shift run-to-run on an unchanged prompt, so a single high score doesn't prove improvement. In a COUNTERPOINT section responding to a16z's Olivia Moore ("we've hit diminishing returns on intelligence"), Every argues that conclusion is only visible if you're measuring the residual work a human still has to do after the model runs — which most companies aren't instrumented to see.

Prescribed action: pick one recurring job, write down five concrete ways the current model gets it wrong with a gold-standard example of each, and test candidate models against those cases — the minimum viable eval.

Curation section

Zero deep-fetches taken this issue: the CentaurBench/Thinkingbox arXiv abstracts and the Yegge/Cursor/Casado pieces are all adequately summarized by Every's own blurbs for RDCO's purposes; none crossed the bar of needing primary-source verification.

Mapping against Ray Data Co

KateBench's specific design — logging not just accept/reject but the residual edits an editor still makes after accepting a suggestion — is a concrete gap check against RDCO's own /verify-* critic family (verify-vault-write, verify-strategic-output, station-critic, design-critic). Those critics already do the right shape of thing (fresh-eyes, PASS/ITERATE/SCRAP, no self-grading), but none of them log an acceptance rate or track residual rework the way KateBench does — so there's currently no way to tell whether the critic gate is catching more over time, plateauing, or just producing PASS-stamped theater. The KateBench caution about a 40-item silent-drop cap inflating the acceptance number is a direct analog to the "false verified stamps" failure mode already logged in feedback_workflow_agent_output_integrity: a critic can look like it's working from a summary metric while quietly missing the tail of what it should have flagged. Concrete next step worth considering: instrument one /verify-* critic (station-critic is the natural pilot, since it already aggregates per-axis verdicts) with a simple accept/reject/still-had-to-fix log, rather than trusting the PASS/ITERATE label as self-evidently improving over time.

The Aaron Levie line — "Enterprises will not be able to go just on vibes" — is also a plain-language restatement of why RDCO runs mechanical audit scripts (e.g. audit-newsletter-outputs.py) alongside LLM-judged critics rather than critics alone: the mechanical layer is the "offline eval" that doesn't drift with a re-run.

⚠️ Sponsorship

Paid third-party sponsor: Attio ("the agentic CRM") — a standard mid-issue paid block with UTM tracking (utm_medium=newsletter_sponsorship), unrelated in subject matter to the evals argument; no bias into the editorial content. Consistent with the prior Every sponsor pattern (Brief, Lightfield, Scribe Optimize — one-off rotating third parties, not a standing relationship). Standard house cross-promo also present: "Want to sponsor Every?" CTA, and the footer bundle pitching Every's own paid tools (Sparkle, Cora, Spiral, Monologue) plus Every All Access — self-promo, not disclosed as sponsorship but worth flagging as the house's own commercial incentive to make AI-tool-building sound universally good ("Everyone's a builder now").

Related