06-reference

every why evals are so hot right now

2026-09-24·reference·source: Every (Context Window)·by Laura Entis
evalsmodel-selectionharness-engineeringagent-frameworkspersonal-benchmarks

"Why Evals Are So Hot Right Now" — Every (Context Window, Laura Entis)

Why this is in the vault

Two independent data points (a hiring-market signal and a usage-growth curve) confirm evals have moved from practitioner niche to mainstream expectation, plus a distinct essay arguing conversational-platform agents (Slack, Teams, WhatsApp) are becoming the next "front door" the way websites once were.

The core argument

Evals — tests of how well an AI system performs a task — are having a moment. Tech influencer Lenny Rachitsky found that of 25 product-management job postings he reviewed last week, nearly half asked for eval-writing experience. At app-monitoring company Sentry, proposed code changes now trigger roughly 1,800 eval runs per month, up from near zero in May, per engineer Ryan Brooks. Senior applied AI engineer Nityesh Agarwal frames the underlying shift: choosing the most intelligent model used to be a safe default, but frontier models are now capable enough that most everyday knowledge work doesn't need full intelligence — so speed, price, and fit with a person's taste matter more than raw capability. Agarwal is helping Every staff build "personal benchmarks" from their own recurring tasks and standards, so a model release can be graded against actual work rather than public leaderboard scores.

Curation section

Mapping against Ray Data Co

The most concrete connection is architectural, not eval-related: RDCO's Channels agent (Mac Mini always-on, iMessage-only per feedback_channels_are_bidirectional and project_channels_agent_setup) is already the exact bet Willie Williams describes — an agent that lives inside a conversational surface the founder already occupies, rather than a dashboard or web app he has to go visit. Every's framing (Slack agents will be as mandatory as websites became) is a validation of that architecture choice, not a new idea to adopt — but it's also a prompt to ask whether RDCO's current single-channel (iMessage-only, Discord retired) posture is under-building the "front door" surface relative to where founders' teams actually work, if RDCO ever needs to deploy an agent for someone other than the founder himself.

On evals specifically: this is the fourth Every piece on personal benchmarks/evals filed to the vault in a month (2026-08-25-every-benchmarks-dont-know-your-job, 2026-09-10-every-evals-for-everyone, 2026-09-21-every-personal-ai-benchmark), so the mapping here is additive rather than novel — Sentry's 1,800-eval-runs/month growth curve (near zero in May) is a useful external data point for how fast the practice scales once adopted, which is the trajectory audit-newsletter-outputs.py and the fresh-eyes critic family (verify-vault-write, verify-strategic-output, verify-dispatch) are already on. The still-unclosed gap flagged in the Sep 10 note stands: RDCO has the eval-as-gate half of this pattern but not the eval-as-model-selector half (feedback_delegation_model_effort_pairing sets model/effort by task-size heuristic, not by benchmarked comparison).

Related

[[2026-09-10-every-evals-for-everyone]] [[2026-08-25-every-benchmarks-dont-know-your-job]] [[project_channels_agent_setup]] [[feedback_delegation_model_effort_pairing]]