"Evals for Everyone" — Every (Context Window, Laura Entis)
Why this is in the vault
Every is building a personal eval for each of its employees — the exact structure RDCO already runs as audit-newsletter-outputs.py, verify-vault-write, and the other fresh-eyes critic gates, just not named or generalized as a standing practice the way Every is now doing it.
The core argument
Public benchmarks (a graduate-level science exam Astra and Fable 5.1 both score 90%+ on) say nothing about whether a model knows where you'd put a comma or how many ideas belong on a slide. Every CEO Dan Shipper's fix: build a personal benchmark per employee — custom evals testing specific parts of that person's work, graded against their own standard for what "good" looks like. Editor in chief Kate Lee's benchmark encodes her copy-editing judgment (comma placement, word choice); head of evals Mike Taylor's encodes a deck rule ("one idea per slide").
The construction method, run by Taylor and applied AI engineer Nityesh Agarwal: collect 5-10 tasks a person regularly hands to AI plus their prompts and source docs, run them across models, then interview the person on what they like/dislike about each output. That feedback becomes pass/fail checks — some programmatic (does the page load, do calculated scores match source data), some subjective (does an AI judge confirm the design avoids a disliked aesthetic). When the human disagrees with the eval's verdict, they don't override it — they tune the checks, because the mismatch means the benchmark is either missing a check or has the wrong definition for one. Once human and AI-judge agree reliably, a single benchmark run tells you which model wins for that task — which is how Taylor discovered a cheaper model (GPT-5.6 Luna) matched Fable and GPT-5.6 Sol on his day-to-day work, freeing him to expand the benchmark toward harder tasks instead of re-running the same comparison by hand.
The piece closes with a five-step DIY workflow: pick a recurring task → save a fixed test case (prompt + source files + first untouched output) → turn your verbal feedback into yes/no checks → have the AI grade the output against those checks and compare its verdict to your own → resolve disagreements by fixing the check wording (not re-grading to get the answer you wanted), then re-validate the revised checklist against a fresh example.
Mapping against Ray Data Co
RDCO already runs this pattern mechanically and didn't name it: audit-newsletter-outputs.py is a literal personal benchmark for vault notes — it encodes the founder's standards (required frontmatter, the 9 newsletter_format values, the ## Related ≥2-vault-link rule, copyright discipline) as pass/fail checks that run after every write, exactly like Taylor's programmatic checks on his dashboard task. The fresh-eyes critic family (verify-vault-write, verify-strategic-output, verify-dispatch) is the subjective-check half — an AI judge scoring against a rubric the way Every's judge scores design against Taylor's dark-mode aversion. And the "turn corrections into checks" workflow is precisely how MEMORY.md's feedback_* entries get created: a founder correction becomes a standing rule an agent is graded against going forward, not a one-off fix. What RDCO is missing that Every has: model selection runs through the benchmark. Taylor's Luna discovery — a cheaper model matching pricier ones on his actual task mix — is the move RDCO hasn't made; feedback_delegation_model_effort_pairing sets model/effort by task-size heuristic, not by a benchmark score, so there's no equivalent to "run the vault-note-writing task across three models, grade against the audit script and a fresh-eyes critic, and let the scores pick the daily driver" the way Every's team visibly rotates models per person in "The daily driver."
Curation section
- Thesis Statements (7 new entries): Every's running collection of contestable claims about the future of human work with AI added seven predictions this week, including Mike Taylor's "Every good idea will have 1,000 clones by morning" and Sahil Lavingia's "Institutions won't need a crisis to change." Tied to Every's inaugural Thesis: 2027 conference, November 5, 2026.
- The daily driver: Every's weekly roundup of which model each team member is using. Notable: head of operations Arielle Shipper dropped Sonnet 5 after a 12-minute doc-comparison task that GPT-5.6 Sol finished in 42 seconds; Mike Taylor called Astra's computer-use "phenomenal" for slide-deck edits but still rates "Claude is still the better writer" after a head-to-head with Astra.
- Straight from Slack: Every is turning its internal Friday show-and-tell into a public series touring how individual team members configure Codex and Claude Code — folder systems, setups — starting with Cora GM Kieran Klaassen.
Related
[[2026-08-25-every-benchmarks-dont-know-your-job]] [[2026-06-08-every-guardrails-review-skills]] [[feedback_delegation_model_effort_pairing]] [[feedback_skills_over_commands]]
⚠️ Sponsorship
No paid third-party sponsor block in this issue. The only promotional content is the standard Every house footer: Every All Access / Builder Pack upsell ($9,000+ in tool credits) and cross-promo for Every's own product bundle (Sparkle, Cora, Spiral, Monologue). Marked sponsor_entity: self — consistent with how other Every Context Window notes in this vault treat the recurring house footer versus the one confirmed one-off paid sponsor (Brief/briefhq.ai).