06-reference

every personal ai benchmark

2026-09-21·reference·source: Every (Also True for Humans)·by Mike Taylor
evalsbenchmarksmodel-selectionharness-engineering

"How to Create Your Own Personal AI Benchmark" — Mike Taylor (Every, Also True for Humans, 2026-09-21)

Why this is in the vault

A working method (not just critique) for building a personal model-selection benchmark, from Every's own head of evals — direct actionable follow-on to the "benchmarks don't know your job" thesis already in the vault (2026-08-25).

The core argument

Public benchmarks (MMLU-Pro, Math Olympiad scores) measure trivia recall, not whether a model can do your job — Wharton's Ethan Mollick's framing, which the author explicitly credits. Taylor, who runs Every's technology consulting practice and is now head of evals, built a personal benchmark instead: a small private set of real work tasks (dashboards, decks, writing) scored consistently across models, so model choice becomes evidence-based rather than vibes.

Method, in order:

  1. Pick 5-10 real recurring tasks from your own work, not synthetic puzzles.
  2. Run each task across candidate models, score results side by side.
  3. Stay above the "discernment horizon" (a term from Steve Yegge) — tasks too easy for a smarter model to show its edge produce flat scores across models; the test has to be hard enough that only some models clear it.
  4. Route trivial/routine tasks to cheap fast models (his example: Luna) and ambitious/open-ended tasks to frontier models (Fable, Sol) — cited an Anthropic finding that model-quality gaps show up mainly on substantial, open-ended tasks, not routine ones.
  5. To scale into a team-wide benchmark: turn each task into a reusable skill, capture brief/output/feedback every time it runs, and accumulate a folder structure (benchmarks/tasks/<task>/<case>/{prompt.txt, context/, gold/, evals.py}) with an index.html for side-by-side review.
  6. Treat the benchmark as perishable — task difficulty has to escalate as models improve, or the benchmark "saturates" (his example: Mollick's "otter on a plane" image test went from hard to trivial and had to be replaced with a much harder prompt).

Mapping against Ray Data Co

Directly names the mechanism RDCO already runs ad hoc — the founder's own model-vibe-checks (the "every-vibe-check-*" series already filed) and Ray's own skill-authoring practice are effectively an unstructured version of exactly this: real-task benchmarking instead of trusting public leaderboards. The benchmarks/tasks/<task>/<case>/{prompt.txt, gold/, evals.py} folder structure is a concrete, adoptable pattern for RDCO's own skill-quality work (parallel to /self-review's rubric-based grading and verify-* critic agents) — worth evaluating whether Ray's skill-improvement loop (/improve) should formalize a gold/ reference-output directory per skill rather than relying on founder spot-checks. The "discernment horizon" concept is also a sharper vocabulary for something RDCO already does implicitly when picking Fable vs Sonnet vs Haiku per task tier (delegation model/effort pairing).

Curation section

No secondary curated items — this issue is a single-argument essay with inline citations (Mollick's "job interview for your AI" piece, Steve Yegge's discernment-horizon framing, Anthropic's Opus/Fable comparison finding). Not a curation or hybrid issue.

⚠️ Sponsorship

No third-party paid sponsor block in this issue — only Every's standard house-promo footer (Every All Access membership + Builder Pack credits, cross-promo for Sparkle/Cora/Spiral/Monologue). No bias risk to the core argument; flagged per standing practice since sponsored always needs a paired entity.

Related

[[2026-08-25-every-benchmarks-dont-know-your-job]] [[2026-09-10-every-evals-for-everyone]] [[feedback_delegation_model_effort_pairing]]