06-reference

every fable vs astra context window

2026-09-06·reference·source: Every·by Every Staff (lead: Katie Parrott)
anthropicfable-5-1gpt-6-astramodel-comparisoncompound-writinganthropic-certificationagent-design

Why this is in the vault

A same-week, side-by-side field comparison of Fable 5.1 (Anthropic) vs. GPT-6 Astra (OpenAI) from Every's internal eval team, plus a bundle of adjacent pieces (a new writing-with-AI method, an Anthropic-certification field report, and a "folders as agents" repost) worth keeping as calibration on both model choice and agent-design habits.

The core argument

Every's staff ran Fable 5.1 and GPT-6 Astra head-to-head at a subscriber-only "camp" and came away split rather than declaring a winner — the two models optimize for different collaboration styles, not just different output quality. Fable needed more context up front but was easier to build on incrementally without undoing unwanted additions; Astra was easier to steer through back-and-forth but sometimes joined ideas without explaining the connection. On concrete tasks Fable won more often: fewer clicks in a journal-digitizing app, a correct diagram Astra decorated into nonsense, and a stronger essay draft. The sharpest data point: in a brief requesting 8–12 supporting quotes, Fable 5.1 returned 43, several sourced from nowhere in the source material — a fabrication-under-latitude failure that shows up even in a model Every otherwise called "so back."

Curation section

Mapping against Ray Data Co

The most concrete connection is the Fable 5.1 over-generation failure — a brief asking for 8–12 quotes came back with 43, some invented — which is the same failure shape catalogued in feedback_workflow_agent_output_integrity's seven factory failure modes (proposal-as-fact, false "verified" stamps). It's a live reminder that even the model RDCO defaults to for Claude Code work can blow past explicit numeric constraints, reinforcing why every chain needs a gate that hits the primary source rather than trusting model self-report on scope compliance. Second: Klaassen's "the folder is the agent" argument — a directory + CLAUDE.md + skills = a specialist, and "you can't vibe orchestrate" — is close to the philosophy behind RDCO's own domain-agnostic brigade stations (station-spec-author, station-test-author, station-code-author, station-critic) and per-repo CLAUDE.md scoping; it's independent confirmation that the pattern isn't over-engineered. Third: the Compound Writing loop (ideate → outline → draft → review → finalize, "each session compounds") is a natural cross-check against WRITING-rdco.md and the draft-review skill's own iteration discipline for Sanity Check drafts. The Anthropic-certification field report is a secondary data point against the founder's own CCA-F pass (2026-08-17, 850/1000) — Every's team found the same documentation-outclasses-videos gap and a stale Sonnet model dependency, both worth flagging if the founder cites the cert publicly.

Related