"Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice" — Dan Shipper and Katie Parrott (Every, Jul 24 2026)
Why this is in the vault
Every's Vibe Check series is RDCO's most reliable signal on frontier model behavior changes — same authors, same testing methodology, same publication that called Opus 4.7's instruction-following regression day-zero and Fable 5's coding ceiling accurately. This edition delivers a finding that directly affects any model-migration decision for the RDCO harness: accumulated system-prompt scaffolding built for prior Claude models actively degrades Opus 5's performance, and stripping it out can reverse that degradation dramatically.
The article evaluates Opus 5 across coding, writing, knowledge work, and agents. Top-line verdict: brilliant in flashes, not the easiest drop-in. Specific findings from the preview:
- Opus 5 fights process-heavy instructions. In Every's first week, it argued with instructions, quit before finishing work, and resisted skills and plugins built for Opus 4.x and Fable. This is the inverse of Opus 4.7's regression — that model became stricter about following instructions; Opus 5 appears to have its own opinions about process and will push back.
- Stripping the scaffolding helped. When Every deleted their accumulated system-prompt layers, Opus 5 improved dramatically: stronger software builds, multi-hour debugging runs without stopping, more rigorous knowledge work output. The reversal is striking — less instruction, more capability.
- Thinking-effort level matters. Low and medium effort settings surfaced more Opus 5 strengths in coding and reduced the "annoying quirks" (insubordination, early stopping). Max thinking effort may cause Opus 5 to over-reason in ways that manifest as instruction-fighting.
- Comparative position: Opus 5 doesn't reach Fable's ceiling. It's also less easy to use day-to-day than GPT-5.6 Sol. It occupies an awkward middle — not the peak performer, not the smoothest drop-in. Its best work requires rebuilding workflows already calibrated to other models.
The full per-category breakdown is subscriber-only; this note is based on the email preview.
⚠️ Sponsorship
Sponsor: Scribe Optimize (scribe.com/optimize). Paid placement — they capture how organizations actually work to direct automation spend. The sponsor is in the process-mining / workflow-discovery space, which happens to be thematically adjacent to the article (auditing workflows before automating them). Treat any implied endorsement of workflow-audit tooling as sponsored. Every also promotes its own products (All Access, Sparkle, Cora, Spiral, Monologue) — standard house upsell, not a paid third-party placement.
Mapping against Ray Data Co
The harness-accumulation risk is the direct hit. RDCO's CLAUDE.md and skills directory have been built and refined across multiple Claude model generations (Opus 4.7 → Fable 5 → Sonnet 4.6). Each model transition has left sediment — instructions, patterns, and tool scaffolding calibrated to prior model behavior. Every's finding says Opus 5 is particularly sensitive to this sediment: it fights it rather than ignoring it or tolerating it. If RDCO ever evaluates Opus 5 for the always-on agent, the right test protocol is not to run it on the current harness and measure degradation — it's to run a stripped harness and measure uplift.
Secondary implications:
- Thinking-effort config matters. If Opus 5 trials happen, start at low/medium effort for coding tasks rather than max. The current harness defaults may not be the right starting point.
- Strategic confirmation: If RDCO is already on Fable 5 for the always-on agent, this article argues against a downgrade to Opus 5. Fable has the higher ceiling and better day-to-day usability per Every's comparative frame. No action needed if Fable is current.
- Model-migration SOP gap: Every's experience surfaces a recurring pattern — each major Claude model transition requires a workflow audit, not just a model swap. This should be a standing SOP item, not an ad-hoc discovery.
Related
- [[2026-06-09-every-vibe-check-fable-5-best-coding-model]] — directly referenced; established Fable 5's ceiling that Opus 5 is now benchmarked against; same authors, same methodology
- [[2026-04-17-every-vibe-check-opus-4-7]] — same series; Opus 4.7 was the last major model-behavior-change finding from Every; established the pattern of "new model breaks old workflows" that Opus 5 repeats with different failure modes
- [[2026-04-15-thariq-claude-code-session-management-1m-context]] — Thariq's session-management guidance directly addresses accumulated context weight; this article reinforces the same concern from the prompt-layer angle (system-prompt sediment, not just context size)