Stop Coding and Start Planning
Re-read at full length 2026-10-03 (paid).
Why this is in the vault
Klaassen's "three fidelities" framework is the clearest rule we have for how much planning a piece of agent work deserves, and his claim that reviewed plans (not code) are what teach the system is the mechanism behind compound engineering. The earlier version of this note was a roughly 200-word preview; this rewrite covers the full essay, including the three-prototype case study and the "50 plan reviews" learning curve. Every resurfaced the piece during a July 2026 offsite.
The core argument
AI made developers sloppy because it made planning feel optional. Klaassen admits to prompting "make this feature work," then spending three hours debugging what a 10-minute research-and-outline session would have prevented. Worse, each vibe-coded feature starts from zero: the system learns nothing.
The contrast he draws is between two prompts. Vibe coding: "Add email validation to the signup form." Planning: ask the agent to research how validation is handled elsewhere in the codebase, check whether the email library already validates, look up good error-message practice, and return three approaches with tradeoffs. One ships a feature; the other ships a feature and "teaches the system how you think."
Case study: five Figma screens in a weekend
For Cora's email bankruptcy launch (a free service that clears a user's inbox without deleting anything important), Klaassen had five designed screens and a weekend. Instead of the old loop of Figma MCP output plus manual pixel-fixing, he spent about one hour building two agents:
- A planning agent that takes a Figma screenshot and writes an implementation plan grounded in the codebase's own components and patterns, stored in GitHub.
- A review agent that screenshots the built page with Puppeteer, diffs it against the design, and iterates until they match.
Because the plan settled what was being built, the reviewer could focus on execution. Result: five screens matching the designs, plus mobile layouts nobody had designed. The workflow is now reusable for the next interface.
The three fidelities
Not all work deserves the same planning budget.
- Fidelity One, the quick fix. One-line changes, copy edits, obvious bugs. Planning is a few lines: reproduce, confirm the fix location, check for similar instances. The category grows with each model: with Sonnet 4.5 he includes codebase-wide pricing changes, dependency migrations and resolving clear PR comments. One example: he gave the agent an error message and "Go fix it."
- Fidelity Two, the sweet spot. Multi-file features with clear scope but non-obvious implementation: moving inline work to a background job, adding an assistant tool call, chasing a bug whose cause is unclear. Planning pays most here because the agent could go off the rails without guidance but executes reliably once guided. Example: "archive by query" for Cora. Parallel research agents (10-20 minutes) found a reusable search-interpreting tool and Gmail's bulk-operation quotas. Without that research it would have passed tests and failed on users archiving thousands of emails.
- Fidelity Three, the big uncertain. Multi-account support, re-architecture, complex integrations, where "done" is unknown. Planning alone cannot remove that uncertainty. He uses vibe planning: build disposable prototypes in a separate environment, learn what breaks, throw them away, then plan. "The prototype is disposable; the knowledge isn't."
Fidelity can change mid-task. Email bankruptcy (processing, say, 53,000 emails) looked like Fidelity Two until research hit Gmail rate limits and timeouts. He built three prototypes of rising complexity: real-time API calls (choked at 1,000 emails), a simple cache (race conditions), and a full queue (the only one that worked). Within a week that settled the architecture. He then split the feature into sequenced Fidelity Two pieces (cache layer, then queue, then API optimization; separately the marketing design and agentic flow), each with its own success criteria. The goal for Fidelity Three is always to break it down into Fidelity Two pieces.
Why plans compound and code does not
Code teaches "how to solve this problem"; planning teaches "how to think about problems like this." When you react to a plan ("too complex," "sequence this differently"), that feedback can become permanent agent instructions. Example: the Figma agent used plain HTML; he corrected it to View Components once, wrote that into the agent's instructions, and it now defaults correctly.
The learning curve he reports: in week one, plans were over-engineered, missed existing patterns and forgot security checks. Three months and more than 50 plan reviews later, plans mostly match how he would approach the problem. He credits accumulated instructions, not better prompting. Better models raise everyone's baseline; your own system improves because it holds your preferences, domain research strategies and known blind spots.
Mapping against Ray Data Co
The most concrete connection is RDCO's plan → tests → implementation rule and the station brigade (spec author, test author, code author, critic). Klaassen supplies the missing piece: a sizing rule. RDCO currently tends to apply full ceremony or none; the three fidelities give a principled default.
- Fidelity Three equals the "structure before research dispatch" lesson. Klaassen's mid-task reclassification and his split into Fidelity Two pieces match the feedback memory that many-entity work needs an index and phases before fan-out.
- "50 plan reviews" is the review-bandwidth argument in miniature. Ben's review time compounds only when each correction is written into a skill or memory file. Corrections given in iMessage and never codified are spent once. Ray's feedback memory files are this mechanism; the gap is making codification automatic after every correction.
- Planner/reviewer split mirrors producer/fresh-eyes critic. The screenshot-diff reviewer is a behavior check against a spec, close to the behavior-critic skill.
- Factory relevance. In the copilot agent factory, most new client skills are Fidelity Two. A research-first plan stage (existing skills, data constraints, quota and governance limits on Snowflake) is where the Gmail-quota class of failure gets caught.
Why it matters for RDCO / The Denominator
- Do: add a fidelity tag (F1/F2/F3) to dispatch prompts and Notion tasks. F1 runs straight through; F2 requires a research-first plan; F3 requires throwaway prototypes before any plan, then decomposition into F2 tickets.
- Do: after any founder correction on a plan, write it into the relevant skill the same turn and log it, so Ray's equivalent of the "50 plan reviews" curve is measurable.
- Write: a Denominator angle: successful enterprise agents fail on platform limits that research would have found (quotas, rate limits, permissions), not on model quality. Klaassen's archive-by-query and 1,000-email choke are clean illustrative cases. Cite as Every's examples.
⚠️ Sponsorship
House content: Cora, Every's own product, is the case study throughout, and the footer promotes Spiral, Sparkle, Cora, Monologue, the Compound Engineering plugin and Every's AI consulting. Bias implication: the success stories (pixel-matching screens, the planning curve) are self-reported by the product's general manager with no failure cases. The framework itself does not depend on any Every product.
Related
- [[2026-02-09-every-compound-engineering-guide]] - the full loop; this essay expands its Plan step.
- [[2026-04-04-compound-engineering]] - earlier vault summary of the methodology.
- [[2026-08-04-every-design-layer-spec-first-ai-agents]] - spec-first agent work from Every, mapped to RDCO's brigade.
- [[2026-03-13-every-compound-engineering-camp]] - the brainstorm and plan steps taught hands-on.
- [[2026-06-24-stratechery-ben-thompson-vibe-coding-adventure]] - outsider view of vibe coding's limits.
- [[2026-01-26-every-claude-code-shipping]] - Klaassen's earlier, pre-planning workflow.
- [[feedback_plan_tests_implementation_order]]