06-reference

every evals for everyone

2026-09-10·reference·source: Every (Context Window)·by Laura Entis
evalspersonal-benchmarksharness-engineeringskills-as-checksagent-model-selection

"Evals for Everyone" — Every (Context Window, Laura Entis)

Why this is in the vault

Every is building a personal eval for each of its employees — the exact structure RDCO already runs as audit-newsletter-outputs.py, verify-vault-write, and the other fresh-eyes critic gates, just not named or generalized as a standing practice the way Every is now doing it.

The core argument

Public benchmarks (a graduate-level science exam Astra and Fable 5.1 both score 90%+ on) say nothing about whether a model knows where you'd put a comma or how many ideas belong on a slide. Every CEO Dan Shipper's fix: build a personal benchmark per employee — custom evals testing specific parts of that person's work, graded against their own standard for what "good" looks like. Editor in chief Kate Lee's benchmark encodes her copy-editing judgment (comma placement, word choice); head of evals Mike Taylor's encodes a deck rule ("one idea per slide").

The construction method, run by Taylor and applied AI engineer Nityesh Agarwal: collect 5-10 tasks a person regularly hands to AI plus their prompts and source docs, run them across models, then interview the person on what they like/dislike about each output. That feedback becomes pass/fail checks — some programmatic (does the page load, do calculated scores match source data), some subjective (does an AI judge confirm the design avoids a disliked aesthetic). When the human disagrees with the eval's verdict, they don't override it — they tune the checks, because the mismatch means the benchmark is either missing a check or has the wrong definition for one. Once human and AI-judge agree reliably, a single benchmark run tells you which model wins for that task — which is how Taylor discovered a cheaper model (GPT-5.6 Luna) matched Fable and GPT-5.6 Sol on his day-to-day work, freeing him to expand the benchmark toward harder tasks instead of re-running the same comparison by hand.

The piece closes with a five-step DIY workflow: pick a recurring task → save a fixed test case (prompt + source files + first untouched output) → turn your verbal feedback into yes/no checks → have the AI grade the output against those checks and compare its verdict to your own → resolve disagreements by fixing the check wording (not re-grading to get the answer you wanted), then re-validate the revised checklist against a fresh example.

Mapping against Ray Data Co

RDCO already runs this pattern mechanically and didn't name it: audit-newsletter-outputs.py is a literal personal benchmark for vault notes — it encodes the founder's standards (required frontmatter, the 9 newsletter_format values, the ## Related ≥2-vault-link rule, copyright discipline) as pass/fail checks that run after every write, exactly like Taylor's programmatic checks on his dashboard task. The fresh-eyes critic family (verify-vault-write, verify-strategic-output, verify-dispatch) is the subjective-check half — an AI judge scoring against a rubric the way Every's judge scores design against Taylor's dark-mode aversion. And the "turn corrections into checks" workflow is precisely how MEMORY.md's feedback_* entries get created: a founder correction becomes a standing rule an agent is graded against going forward, not a one-off fix. What RDCO is missing that Every has: model selection runs through the benchmark. Taylor's Luna discovery — a cheaper model matching pricier ones on his actual task mix — is the move RDCO hasn't made; feedback_delegation_model_effort_pairing sets model/effort by task-size heuristic, not by a benchmark score, so there's no equivalent to "run the vault-note-writing task across three models, grade against the audit script and a fresh-eyes critic, and let the scores pick the daily driver" the way Every's team visibly rotates models per person in "The daily driver."

Curation section

Related

[[2026-08-25-every-benchmarks-dont-know-your-job]] [[2026-06-08-every-guardrails-review-skills]] [[feedback_delegation_model_effort_pairing]] [[feedback_skills_over_commands]]

⚠️ Sponsorship

No paid third-party sponsor block in this issue. The only promotional content is the standard Every house footer: Every All Access / Builder Pack upsell ($9,000+ in tool credits) and cross-promo for Every's own product bundle (Sparkle, Cora, Spiral, Monologue). Marked sponsor_entity: self — consistent with how other Every Context Window notes in this vault treat the recurring house footer versus the one confirmed one-off paid sponsor (Brief/briefhq.ai).