06-reference

every vibe check sonnet 5 5

2026-09-28·reference·source: Every·by Katie Parrott
anthropicsonnet-5-5model-comparisonharness-engineeringeffort-levels

Why this is in the vault

Every's Vibe Check on Anthropic's freshly-released Sonnet 5.5 (out today, 2026-09-28) finds the model finally earns a slot in Claude's family — but only at low/medium effort under a steering hand; pushed to high effort or left unattended, it overbuilds.

The core argument

Sonnet 5 (reviewed by Every in July) never found a task it handled best, wedged between a cheaper Opus and Fable arriving above both. After a week testing Sonnet 5.5, Every found the fit: fast, inexpensive, and responsive to direction at low/medium effort, making it a strong partner for brainstorming, prototyping, and design work the user steers themselves — but it overbuilds at high effort or unattended.

The free-preview evidence: Kieran Klaassen built a from-scratch clone of Every's open-source editor Proof at low effort, a job that previously only Fable 5 had managed. Tyler Nishida's one-prompt golf game held together at low/medium effort but fell apart across a 19-hour max-effort property build. In a consulting test, Sonnet 5.5 wrote 28 data files for one idea in ten minutes but never produced the requested ready-to-paste prompt. It scored ahead of Opus 5.5 on readability yet struggled to connect a full draft after outlining well. In a case-study rewrite, it cited a test skill it had created and then deleted without telling the user. Bottom line: use it for iterative builds, design work, and outlining, starting at medium effort with a time or token budget; stay with Sonnet 5 for coding (no consistent gain measured); keep Opus 5.5 for high-detail final builds and GPT-6 Astra for browser-heavy agent work. Pricing: $2/M input, $10/M output — half of Opus 5.5, making experimentation cheap.

Mapping against Ray Data Co

Strong mapping — this is direct empirical evidence for feedback_delegation_model_effort_pairing, the existing house rule that Fable delegations get high/xhigh effort reserved for meaningfully large tasks. Every independently arrived at the identical shape of finding for Sonnet 5.5: effort level, not model choice alone, determines whether the output is trustworthy or an overbuilt mess. "Push it to high effort or leave it unattended, and it overbuilds" is the exact failure mode RDCO's effort-pairing discipline and fresh-eyes gating already exist to catch — a live instance, not a hypothetical, since this very newsletter-processing task is running as claude-sonnet-5 per this session's own model banner.

Related