06-reference

every vibe check opus 5 5 codex converts

2026-09-22·reference·source: Every·by Katie Parrott
anthropicopus-5-5fable-5-1codexmodel-comparisonharness-engineering

Why this is in the vault

Every's Vibe Check on Anthropic's Opus 5.5 — pitched at roughly Fable 5.1-level performance for about 60% less per token — with the notable finding that it's pulling builders who'd switched to Codex back to Claude, while still failing as a reliable finisher under time or precision pressure.

The core argument

Every tested Opus 5.5 across coding, design, writing, and consulting tasks and found the price/performance claim "surprisingly credible." Kieran Klaassen made it his daily driver in place of Fable 5.1; former Claude users Mike Taylor and Tyler Nishida, who'd both moved to Codex, are reconsidering. The verdict isn't unanimous — Dan Shipper still prefers Codex for knowledge work and writing, and the piece is explicit that Opus remains an unreliable finisher when time or precision is tight.

The free-preview anecdotes (full comparison behind Every's paywall) anchor both sides of that split. On the capability side: Opus's Ruby code handled 427 requests per second and met 17 of 20 latency budgets, beating GPT-6 Astra on that measure; Mike's reviewable 251-line Rails patch passed four of five automated checks; a one-prompt golf game ran for nearly two hours; a negotiation task reached agreement after 12 rounds; and Opus produced the most readable prose Every has measured. On the failure side: a separate app burned 5.9 million tokens before its core screens threw errors; a timed training-design task produced handouts but never delivered the requested schedule; and despite readable prose, Opus buried ideas that Astra reliably surfaced first. Every's stated policy: reach for Opus 5.5 on visual products and creative collaboration, keep Fable for problems too large to inspect easily, keep Astra or GPT-5.6 Sol nearby when the deliverable has a clock — and set Opus a budget and a stopping point regardless. Disclosure line in the email: "Anthropic provided Every with pre-launch access and had no input on the review."

Mapping against Ray Data Co

Strong mapping — the 5.9-million-token runaway-before-failure example is a direct, named instance of the exact risk RDCO's fresh-eyes critic gates and cost-discipline rules exist to catch, not a hypothetical.

Related