"Benchmarks Don't Know Your Job" — Katie Parrott (Every, Context Window, 2026-08-25)
Why this is in the vault
Sharpest public articulation this month of the "leaderboard score ≠ your task" problem, illustrated with Every's own internal eval (KateBench) rather than a vendor pitch.
The core argument
Companies that spend heavily on AI models often know their spend and public benchmark rankings but not whether a model actually saves time or produces trustworthy work on their specific tasks. Mercor CEO Brendan Foody and Box CEO Aaron Levie ("Enterprises will not be able to go just on vibes") are cited making the same point: public benchmarks show relative capability, not whether a model caught the clause your lawyers care about or preserved your house style. With model choice multiplying (open-source token share on Vercel's platform going 28%→62% in two months; legal-worker Codex adoption up 108x since February per a16z), picking the top-scoring model is no longer a real strategy.
Every's answer is task-specific "offline evals" built from real work. Its flagship internal instrument, KateBench, is an AI copyeditor trained on ~30,000 of editor-in-chief Kate Lee's past edits; it suggests changes in a Google Doc and tracks which ones editors accept, reject, or still rewrite afterward. The piece is candid about the instrument's own failure modes: KateBench's engineer, Jannik Jung, found the tool silently dropped suggestions past a 40-item cap (raised to 80) — meaning a headline 85–90% acceptance rate was partly an artifact of undercounting, not quality — and that acceptance rates shift run-to-run on an unchanged prompt, so a single high score doesn't prove improvement. In a COUNTERPOINT section responding to a16z's Olivia Moore ("we've hit diminishing returns on intelligence"), Every argues that conclusion is only visible if you're measuring the residual work a human still has to do after the model runs — which most companies aren't instrumented to see.
Prescribed action: pick one recurring job, write down five concrete ways the current model gets it wrong with a gold-standard example of each, and test candidate models against those cases — the minimum viable eval.
Curation section
- Fences, not sandboxes (Steve Yegge, yegge.ai) — argues coding agents need explicit behavioral rules ("fences") rather than only technical containment ("sandboxes").
- Git at any scale (Cursor engineering blog) — Cursor rebuilt its Git hosting because coding agents spin up huge numbers of short-lived repos; new system stores every change in the cloud and materializes/discards working copies on demand.
- Fable's data-retention gap (Martin Casado, a16z, via X) — argues Fable's low enterprise share is a privacy problem: no zero-data-retention option, a nonstarter under strict data policies.
- CentaurBench (arXiv 2608.18554) — the model best at solo task completion was not necessarily the best helper; on 5 of 7 tasks a different model was better at improving a weaker model's first attempt.
- Thinkingbox (arXiv 2608.19741) — the strongest coding model passed 65% of single attempts but dropped to 25% success when required to perform reliably across 20 consecutive attempts.
- Also in this issue (Every-original, not a third-party link): "The Dryer Has an Agent Team" — Every designer Tyler Nishida used Grok Bot/Grok Build to spin up a six-agent crew (via Raspberry Pi) monitoring his off-grid solar system in Hawaii, to answer whether there's enough power to run the dryer. Framed as a lighthearted closer, not analysis — no deep-fetch warranted.
Zero deep-fetches taken this issue: the CentaurBench/Thinkingbox arXiv abstracts and the Yegge/Cursor/Casado pieces are all adequately summarized by Every's own blurbs for RDCO's purposes; none crossed the bar of needing primary-source verification.
Mapping against Ray Data Co
KateBench's specific design — logging not just accept/reject but the residual edits an editor still makes after accepting a suggestion — is a concrete gap check against RDCO's own /verify-* critic family (verify-vault-write, verify-strategic-output, station-critic, design-critic). Those critics already do the right shape of thing (fresh-eyes, PASS/ITERATE/SCRAP, no self-grading), but none of them log an acceptance rate or track residual rework the way KateBench does — so there's currently no way to tell whether the critic gate is catching more over time, plateauing, or just producing PASS-stamped theater. The KateBench caution about a 40-item silent-drop cap inflating the acceptance number is a direct analog to the "false verified stamps" failure mode already logged in feedback_workflow_agent_output_integrity: a critic can look like it's working from a summary metric while quietly missing the tail of what it should have flagged. Concrete next step worth considering: instrument one /verify-* critic (station-critic is the natural pilot, since it already aggregates per-axis verdicts) with a simple accept/reject/still-had-to-fix log, rather than trusting the PASS/ITERATE label as self-evidently improving over time.
The Aaron Levie line — "Enterprises will not be able to go just on vibes" — is also a plain-language restatement of why RDCO runs mechanical audit scripts (e.g. audit-newsletter-outputs.py) alongside LLM-judged critics rather than critics alone: the mechanical layer is the "offline eval" that doesn't drift with a re-run.
⚠️ Sponsorship
Paid third-party sponsor: Attio ("the agentic CRM") — a standard mid-issue paid block with UTM tracking (utm_medium=newsletter_sponsorship), unrelated in subject matter to the evals argument; no bias into the editorial content. Consistent with the prior Every sponsor pattern (Brief, Lightfield, Scribe Optimize — one-off rotating third parties, not a standing relationship). Standard house cross-promo also present: "Want to sponsor Every?" CTA, and the footer bundle pitching Every's own paid tools (Sparkle, Cora, Spiral, Monologue) plus Every All Access — self-promo, not disclosed as sponsorship but worth flagging as the house's own commercial incentive to make AI-tool-building sound universally good ("Everyone's a builder now").
Related
- [[2026-04-08-better-harness-evals-hill-climbing]] — the vault's existing harness-evals thesis; this issue is a concrete, single-company case study of the same "hill-climbing on your own eval, not the public leaderboard" argument.
- [[2026-05-28-every-vibe-check-opus-4-8]] — same sender/format precedent (Every Vibe Check + rotating third-party sponsor pattern), useful for cross-referencing Every's sponsor rotation.
- [[feedback_workflow_agent_output_integrity]] — the "false verified stamps" failure mode this issue's KateBench caution directly parallels.