06-reference

every vibe check opus5

2026-07-24·reference·source: Every·by Dan Shipper and Katie Parrott

"Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice" — Dan Shipper and Katie Parrott (Every, Jul 24 2026)

Why this is in the vault

Every's Vibe Check series is RDCO's most reliable signal on frontier model behavior changes — same authors, same testing methodology, same publication that called Opus 4.7's instruction-following regression day-zero and Fable 5's coding ceiling accurately. This edition delivers a finding that directly affects any model-migration decision for the RDCO harness: accumulated system-prompt scaffolding built for prior Claude models actively degrades Opus 5's performance, and stripping it out can reverse that degradation dramatically.

The article evaluates Opus 5 across coding, writing, knowledge work, and agents. Top-line verdict: brilliant in flashes, not the easiest drop-in. Specific findings from the preview:

The full per-category breakdown is subscriber-only; this note is based on the email preview.

⚠️ Sponsorship

Sponsor: Scribe Optimize (scribe.com/optimize). Paid placement — they capture how organizations actually work to direct automation spend. The sponsor is in the process-mining / workflow-discovery space, which happens to be thematically adjacent to the article (auditing workflows before automating them). Treat any implied endorsement of workflow-audit tooling as sponsored. Every also promotes its own products (All Access, Sparkle, Cora, Spiral, Monologue) — standard house upsell, not a paid third-party placement.

Mapping against Ray Data Co

The harness-accumulation risk is the direct hit. RDCO's CLAUDE.md and skills directory have been built and refined across multiple Claude model generations (Opus 4.7 → Fable 5 → Sonnet 4.6). Each model transition has left sediment — instructions, patterns, and tool scaffolding calibrated to prior model behavior. Every's finding says Opus 5 is particularly sensitive to this sediment: it fights it rather than ignoring it or tolerating it. If RDCO ever evaluates Opus 5 for the always-on agent, the right test protocol is not to run it on the current harness and measure degradation — it's to run a stripped harness and measure uplift.

Secondary implications:

Related