"How to Create Your Own Personal AI Benchmark" — Mike Taylor (Every, Also True for Humans, 2026-09-21)
Why this is in the vault
A working method (not just critique) for building a personal model-selection benchmark, from Every's own head of evals — direct actionable follow-on to the "benchmarks don't know your job" thesis already in the vault (2026-08-25).
The core argument
Public benchmarks (MMLU-Pro, Math Olympiad scores) measure trivia recall, not whether a model can do your job — Wharton's Ethan Mollick's framing, which the author explicitly credits. Taylor, who runs Every's technology consulting practice and is now head of evals, built a personal benchmark instead: a small private set of real work tasks (dashboards, decks, writing) scored consistently across models, so model choice becomes evidence-based rather than vibes.
Method, in order:
- Pick 5-10 real recurring tasks from your own work, not synthetic puzzles.
- Run each task across candidate models, score results side by side.
- Stay above the "discernment horizon" (a term from Steve Yegge) — tasks too easy for a smarter model to show its edge produce flat scores across models; the test has to be hard enough that only some models clear it.
- Route trivial/routine tasks to cheap fast models (his example: Luna) and ambitious/open-ended tasks to frontier models (Fable, Sol) — cited an Anthropic finding that model-quality gaps show up mainly on substantial, open-ended tasks, not routine ones.
- To scale into a team-wide benchmark: turn each task into a reusable skill, capture brief/output/feedback every time it runs, and accumulate a folder structure (
benchmarks/tasks/<task>/<case>/{prompt.txt, context/, gold/, evals.py}) with anindex.htmlfor side-by-side review. - Treat the benchmark as perishable — task difficulty has to escalate as models improve, or the benchmark "saturates" (his example: Mollick's "otter on a plane" image test went from hard to trivial and had to be replaced with a much harder prompt).
Mapping against Ray Data Co
Directly names the mechanism RDCO already runs ad hoc — the founder's own model-vibe-checks (the "every-vibe-check-*" series already filed) and Ray's own skill-authoring practice are effectively an unstructured version of exactly this: real-task benchmarking instead of trusting public leaderboards. The benchmarks/tasks/<task>/<case>/{prompt.txt, gold/, evals.py} folder structure is a concrete, adoptable pattern for RDCO's own skill-quality work (parallel to /self-review's rubric-based grading and verify-* critic agents) — worth evaluating whether Ray's skill-improvement loop (/improve) should formalize a gold/ reference-output directory per skill rather than relying on founder spot-checks. The "discernment horizon" concept is also a sharper vocabulary for something RDCO already does implicitly when picking Fable vs Sonnet vs Haiku per task tier (delegation model/effort pairing).
Curation section
No secondary curated items — this issue is a single-argument essay with inline citations (Mollick's "job interview for your AI" piece, Steve Yegge's discernment-horizon framing, Anthropic's Opus/Fable comparison finding). Not a curation or hybrid issue.
⚠️ Sponsorship
No third-party paid sponsor block in this issue — only Every's standard house-promo footer (Every All Access membership + Builder Pack credits, cross-promo for Sparkle/Cora/Spiral/Monologue). No bias risk to the core argument; flagged per standing practice since sponsored always needs a paired entity.
Related
[[2026-08-25-every-benchmarks-dont-know-your-job]] [[2026-09-10-every-evals-for-everyone]] [[feedback_delegation_model_effort_pairing]]