Why this is in the vault
A same-week, side-by-side field comparison of Fable 5.1 (Anthropic) vs. GPT-6 Astra (OpenAI) from Every's internal eval team, plus a bundle of adjacent pieces (a new writing-with-AI method, an Anthropic-certification field report, and a "folders as agents" repost) worth keeping as calibration on both model choice and agent-design habits.
The core argument
Every's staff ran Fable 5.1 and GPT-6 Astra head-to-head at a subscriber-only "camp" and came away split rather than declaring a winner — the two models optimize for different collaboration styles, not just different output quality. Fable needed more context up front but was easier to build on incrementally without undoing unwanted additions; Astra was easier to steer through back-and-forth but sometimes joined ideas without explaining the connection. On concrete tasks Fable won more often: fewer clicks in a journal-digitizing app, a correct diagram Astra decorated into nonsense, and a stronger essay draft. The sharpest data point: in a brief requesting 8–12 supporting quotes, Fable 5.1 returned 43, several sourced from nowhere in the source material — a fabrication-under-latitude failure that shows up even in a model Every otherwise called "so back."
Curation section
- "Vibe Check: GPT-6 Astra Is a Big Upgrade With Some Bad Habits" (Katie Parrott) — Astra scored 71/100 on Every's Senior Engineer Bench (up from GPT-5.6's 56) but overbuilds simple tasks and struggles to judge when its own work is done; verdict still favors Fable 5.1 for shipping.
- "Vibe Check: Fable 5.1 — Anthropic Is So Back (Again)" (Katie Parrott, Dan Shipper) — Fable 5.1 matches Opus 5's agent results in ~60% of the time and half the tokens, with prose that reads less "Claudeish," but over-delivers past prompted limits (the 43-vs-8–12-quotes example).
- "Compound Writing: The Ultimate Guide" (Katie Parrott) — Parrott open-sources her two-year AI-writing system: ideate/interview → outline → draft → review → finalize, adapted from Kieran Klaassen's Compound Engineering loop, built so each writing session compounds instead of restarting cold.
- "How an Every Staff Writer Developed Compound Writing" (Kaushik Viswanath, AI & I podcast) — origin story: Parrott built the method starting from a $20/month ChatGPT subscription after a 2023 layoff.
- "What We Learned From 15 Hours of Anthropic Certification Training" (Natalia Quintero) — a third of Every's team took Anthropic's four-course certification; verdict is it's solid "AI 101" vocabulary-alignment (skills, MCPs, APIs) but the documentation outclasses the training videos, and one course still runs on a retired Sonnet model.
- "The Folder Is the Agent" (Kieran Klaassen, republished) — after three months chasing agent "swarms," Klaassen argues a directory with the right context, skills, and a CLAUDE.md IS a specialist agent; "you can't vibe orchestrate" — build and trust the single agent first, then orchestrate.
- Thesis Statements — seven short, contestable predictions on the future of AI-assisted work from outside builders (M13, Writerbuilder, Skyfall AI, Synthesia, Little Plains, Ben's Bites), teasing Every's Nov. 5, 2026 Thesis: 2027 conference.
- Alignment column (Ashwin Sharma) — a drug-discovery study found AI-generated (and factually wrong) molecule descriptions sometimes improved model predictions; the hypothesized mechanism is that hallucinations can nudge a model toward useful counterfactuals rather than just being errors to dismiss.
Mapping against Ray Data Co
The most concrete connection is the Fable 5.1 over-generation failure — a brief asking for 8–12 quotes came back with 43, some invented — which is the same failure shape catalogued in feedback_workflow_agent_output_integrity's seven factory failure modes (proposal-as-fact, false "verified" stamps). It's a live reminder that even the model RDCO defaults to for Claude Code work can blow past explicit numeric constraints, reinforcing why every chain needs a gate that hits the primary source rather than trusting model self-report on scope compliance. Second: Klaassen's "the folder is the agent" argument — a directory + CLAUDE.md + skills = a specialist, and "you can't vibe orchestrate" — is close to the philosophy behind RDCO's own domain-agnostic brigade stations (station-spec-author, station-test-author, station-code-author, station-critic) and per-repo CLAUDE.md scoping; it's independent confirmation that the pattern isn't over-engineered. Third: the Compound Writing loop (ideate → outline → draft → review → finalize, "each session compounds") is a natural cross-check against WRITING-rdco.md and the draft-review skill's own iteration discipline for Sanity Check drafts. The Anthropic-certification field report is a secondary data point against the founder's own CCA-F pass (2026-08-17, 850/1000) — Every's team found the same documentation-outclasses-videos gap and a stale Sonnet model dependency, both worth flagging if the founder cites the cert publicly.
Related
- [[2026-06-09-every-vibe-check-fable-5-best-coding-model]]
- [[2026-02-09-every-compound-engineering-guide]]
- [[2026-08-31-every-anthropic-certification-training]]
- [[project_phdata_cert_escalator_path]]
- [[feedback_workflow_agent_output_integrity]]