Why this is in the vault
Every's head-to-head test of OpenAI's new GPT-6 Astra against Anthropic's Fable 5.1 on real product-building tasks — the team's verdict is that Astra wins on writing, computer-use, and visual design, but Fable 5.1 still has better instincts for the actual job of building a complete product, staying inside scope where Astra over-adds.
The core argument
Every ran GPT-6 Astra — OpenAI's newest and, per the company, most intelligent model — through the same style of test it uses for every major release. Astra produced writing, consulting deliverables, and visual designs the team wanted to keep; Dan Shipper was especially enthusiastic about its computer-use ability to operate software directly through an interface. But the piece's throughline is that a strong first result isn't the same as a finished product: on Every's more complicated app-building tests, Fable 5.1 still outperformed Astra.
The recurring failure mode was scope creep — Astra kept adding copy and interface elements the task never asked for. Three concrete examples anchor the piece (email preview only; the full comparison sits behind Every's paywall): Astra wrote the entire first draft of this review itself in one shot, well enough that Dan initially thought Katie Parrott had written it. Mike Taylor had Astra build an AI-training curriculum from employee interviews that came close to the course he actually wanted to teach. Kieran Klaassen asked for a "cozy island" and got one — plus an unrequested guided-breathing exercise bolted on top. Every's stated takeaway: reach for Astra for writing, consulting work, and visual prototypes; keep reaching for Fable 5.1 for anything that counts as a complicated product.
Mapping against Ray Data Co
Strong mapping — this is third-party, cross-vendor evidence for a scope-discipline risk RDCO's own harness is built to guard against, landing directly on the Anthropic-aligned bet.
- Astra's "adds what wasn't asked for" failure is the exact scope-creep RDCO's dispatch discipline exists to catch. The cozy-island-plus-unrequested-breathing-exercise example is a clean illustration of a model over-delivering past the actual spec — functionally the same failure class as the fabricated-quotes and ignored-stop-request findings in the vault's 2026-09-01 Fable 5.1 vibe-check note (
feedback_workflow_agent_output_integrity's "proposal-as-fact" and unbounded-agency modes). RDCO's answer to that class of risk isn't "trust the model's judgment," it's the fresh-eyes critic loop (verify-vault-write, verify-dispatch, verify-strategic-output) and the "complete, then ask for review — build the whole thing with labeled guesses" rule, which explicitly assumes agents can and will over-scope without a gate. Every's independent finding that Astra does this more than Fable 5.1 is a real (if single-source) data point that switching RDCO's primary build tooling to an OpenAI model would raise, not lower, reliance on that gate. - Corroborates the current Anthropic-aligned tooling bet without being promotional about it. Every runs this same Vibe Check format across both vendors on a regular cadence (GPT-5.6 Sol, Opus 5, Sonnet 5, and now GPT-6 Astra vs. Fable 5.1 all get the same treatment), and this issue carries no third-party sponsor block — just the standard house cross-promo (Sparkle/Cora/Spiral/Monologue, Every All Access) also present in the 2026-09-01 note. That symmetry is worth noting explicitly: Every isn't a captured Anthropic mouthpiece, so "Fable 5.1 still wins for building a product" reads as an independent editorial call rather than vendor favoritism — useful corroboration for RDCO's Anthropic-aligned harness-engineering thesis, but a single outlet's take, not proof.
- Caveat: this note runs on the free-preview only. Every's full comparison (behind the paywall — "every benchmark and our final verdict") likely has more granular per-task detail than the three anecdotes the email surfaced. Treat the "Fable 5.1 wins on complicated apps" verdict as directionally reliable but not yet benchmark-verified from RDCO's side.
Related
- [[2026-09-01-every-fable-5-1-vibe-check]]
- [[2026-07-09-every-vibe-check-gpt56-sol]]
- [[2026-09-02-stratechery-fable-5-1-enterprise-frontier-safeguards]]
- [[feedback_workflow_agent_output_integrity]]
- [[feedback_complete_then_ask_for_review]]