Why this is in the vault
Every's Vibe Check on Anthropic's Opus 5.5 — pitched at roughly Fable 5.1-level performance for about 60% less per token — with the notable finding that it's pulling builders who'd switched to Codex back to Claude, while still failing as a reliable finisher under time or precision pressure.
The core argument
Every tested Opus 5.5 across coding, design, writing, and consulting tasks and found the price/performance claim "surprisingly credible." Kieran Klaassen made it his daily driver in place of Fable 5.1; former Claude users Mike Taylor and Tyler Nishida, who'd both moved to Codex, are reconsidering. The verdict isn't unanimous — Dan Shipper still prefers Codex for knowledge work and writing, and the piece is explicit that Opus remains an unreliable finisher when time or precision is tight.
The free-preview anecdotes (full comparison behind Every's paywall) anchor both sides of that split. On the capability side: Opus's Ruby code handled 427 requests per second and met 17 of 20 latency budgets, beating GPT-6 Astra on that measure; Mike's reviewable 251-line Rails patch passed four of five automated checks; a one-prompt golf game ran for nearly two hours; a negotiation task reached agreement after 12 rounds; and Opus produced the most readable prose Every has measured. On the failure side: a separate app burned 5.9 million tokens before its core screens threw errors; a timed training-design task produced handouts but never delivered the requested schedule; and despite readable prose, Opus buried ideas that Astra reliably surfaced first. Every's stated policy: reach for Opus 5.5 on visual products and creative collaboration, keep Fable for problems too large to inspect easily, keep Astra or GPT-5.6 Sol nearby when the deliverable has a clock — and set Opus a budget and a stopping point regardless. Disclosure line in the email: "Anthropic provided Every with pre-launch access and had no input on the review."
Mapping against Ray Data Co
Strong mapping — the 5.9-million-token runaway-before-failure example is a direct, named instance of the exact risk RDCO's fresh-eyes critic gates and cost-discipline rules exist to catch, not a hypothetical.
- The token-burn failure is the concrete case for "verification belongs to an independent worker." An app that spends 5.9M tokens before its core screens even throw errors is a model that kept working long past the point a human would have stopped it — the same shape of risk
feedback_verification_independent_worker_patternand the fresh-eyes critic skills (verify-strategic-output, verify-vault-write, verify-dispatch) are built to gate: don't let the producing agent be the one who certifies its own output is done and correct. "Set Opus a budget and a stopping point" is Every independently arriving at the same fix RDCO already encodes as a hard rule rather than a suggestion. - Directly relevant to the phData Anthropic cert bet, not just abstract model trivia. RDCO's harness-engineering thesis is explicitly Anthropic-aligned (
project_phdata_cert_escalator_path,project_l5_north_star_strategic_direction) — a credible claim that Opus 5.5 undercuts Fable-tier pricing by ~60% while pulling Codex converts back is evidence the bet on Claude-first tooling is holding up against the OpenAI alternative, not evidence to act on yet (single-source, paywall-truncated, Anthropic gave pre-launch access). - Caveat: this is a free-preview email, not the full Vibe Check. The per-task detail (which checks the Rails patch failed, what the 5.9M-token app actually was) sits behind Every's paywall; treat the "credible price/performance claim" verdict as directionally useful, not benchmark-verified from RDCO's side.
Related
- [[2026-09-01-every-fable-5-1-vibe-check]]
- [[2026-07-24-every-vibe-check-opus5]]
- [[2026-09-03-every-gpt6-astra-vibe-check]]
- [[feedback_verification_independent_worker_pattern]]
- [[project_phdata_cert_escalator_path]]