"Why Evals Are So Hot Right Now" — Every (Context Window, Laura Entis)
Why this is in the vault
Two independent data points (a hiring-market signal and a usage-growth curve) confirm evals have moved from practitioner niche to mainstream expectation, plus a distinct essay arguing conversational-platform agents (Slack, Teams, WhatsApp) are becoming the next "front door" the way websites once were.
The core argument
Evals — tests of how well an AI system performs a task — are having a moment. Tech influencer Lenny Rachitsky found that of 25 product-management job postings he reviewed last week, nearly half asked for eval-writing experience. At app-monitoring company Sentry, proposed code changes now trigger roughly 1,800 eval runs per month, up from near zero in May, per engineer Ryan Brooks. Senior applied AI engineer Nityesh Agarwal frames the underlying shift: choosing the most intelligent model used to be a safe default, but frontier models are now capable enough that most everyday knowledge work doesn't need full intelligence — so speed, price, and fit with a person's taste matter more than raw capability. Agarwal is helping Every staff build "personal benchmarks" from their own recurring tasks and standards, so a model release can be graded against actual work rather than public leaderboard scores.
Curation section
- Mini-Vibe Check — Grok 4.7 is uneven: Every's Grok 4.6 power users found the 4.7 upgrade unusable pre-launch, then split post-launch. Designer Tyler Nishida (97% of his tracked code events were Grok last month) came around after the model's live performance improved and he needed less micromanagement; Cora GM Kieran Klaassen stayed negative ("not even close" to Opus 5.5). Staff writer Katie Parrott's writing tests were uniformly bad — the model produced Elon-Musk-flavored declarative prose with "no cadence or rhythm." Verdict: fine for coding, skip it for writing.
- Steal This Workflow — turning articles into short videos: head of social Becky Isjwara's 4-step process for viral clips: (1) have Claude propose an opening hook and scene sequence from the article URL, (2) storyboard each scene as a still image using the cover's color palette, (3) render via Hyperframes (Fable, in Claude desktop) plus Flora/MiniMax for illustration animation, (4) iterate against Claude's built-in video preview, then save the corrected workflow as a reusable skill.
- Thesis Statements — five new entries: contestable claims about the future of human work with AI, from Drew Breunig (Cmpnd), Sean Campbell (George Fox MBA), Geoffrey Litt (Notion), Daniela Rus (MIT CSAIL), and Arielle Shipper (Every) — tied to Every's inaugural Thesis: 2027 conference, November 5, 2026.
- Jagged Frontier — Slack agents are the new websites: Willie Williams argues conversational platforms (Slack, Teams, Google Chat, WhatsApp, Telegram) are becoming the next mandatory front door, the way websites went from optional to essential for businesses over a decade. Agents deployed there use context already present in the platform instead of requiring a separate destination; Every is building its own Slack-native archive agent.
Mapping against Ray Data Co
The most concrete connection is architectural, not eval-related: RDCO's Channels agent (Mac Mini always-on, iMessage-only per feedback_channels_are_bidirectional and project_channels_agent_setup) is already the exact bet Willie Williams describes — an agent that lives inside a conversational surface the founder already occupies, rather than a dashboard or web app he has to go visit. Every's framing (Slack agents will be as mandatory as websites became) is a validation of that architecture choice, not a new idea to adopt — but it's also a prompt to ask whether RDCO's current single-channel (iMessage-only, Discord retired) posture is under-building the "front door" surface relative to where founders' teams actually work, if RDCO ever needs to deploy an agent for someone other than the founder himself.
On evals specifically: this is the fourth Every piece on personal benchmarks/evals filed to the vault in a month (2026-08-25-every-benchmarks-dont-know-your-job, 2026-09-10-every-evals-for-everyone, 2026-09-21-every-personal-ai-benchmark), so the mapping here is additive rather than novel — Sentry's 1,800-eval-runs/month growth curve (near zero in May) is a useful external data point for how fast the practice scales once adopted, which is the trajectory audit-newsletter-outputs.py and the fresh-eyes critic family (verify-vault-write, verify-strategic-output, verify-dispatch) are already on. The still-unclosed gap flagged in the Sep 10 note stands: RDCO has the eval-as-gate half of this pattern but not the eval-as-model-selector half (feedback_delegation_model_effort_pairing sets model/effort by task-size heuristic, not by benchmarked comparison).
Related
[[2026-09-10-every-evals-for-everyone]] [[2026-08-25-every-benchmarks-dont-know-your-job]] [[project_channels_agent_setup]] [[feedback_delegation_model_effort_pairing]]