Compound Engineering: The Definitive Guide
Re-read at full length 2026-10-03 (paid).
Why this is in the vault
This is Every's canonical reference for the compound engineering loop (Plan, Work, Review, Compound), its plugin, a five-stage adoption ladder, and a set of team norms for agent-written code. It is the most complete external spec of a practice RDCO already runs, so it is useful as a checklist and as citable language. The earlier version of this note was a roughly 200-word teaser built from the guide's intro; this rewrite covers the full guide.
The core argument
Klaassen's thesis: each unit of engineering work should make the next unit easier. Most codebases go the other way. Every feature adds complexity, and after a decade teams spend more time negotiating with old code than building new code. Compound engineering reverses this by turning bug fixes, patterns and review findings into codified knowledge that the agent reads next time. Every says it runs five products (Cora, Monologue, Sparkle, Spiral, Every.to) mostly with one-person engineering teams on this system.
The main loop
Plan → Work → Review → Compound → Repeat. The first three steps are ordinary engineering. The fourth step is the differentiator; skip it and "you've done traditional engineering with AI assistance."
- Plan: understand the requirement, research the codebase, research external docs and best practices, design the solution, validate the plan.
- Work: isolate in a git worktree or branch, let the agent execute step by step, run tests, linters and type checks after each change, track progress, adapt the plan when something breaks. If you trust the plan, you do not watch every line.
- Review: several specialized reviewer agents in parallel; findings ranked P1 (must fix), P2 (should fix), P3 (nice to fix); the agent resolves them; fixes are validated; the failure pattern is written down.
- Compound: capture what worked and what did not, tag it with YAML frontmatter so it is findable, push the pattern into CLAUDE.md or a new agent, then ask "would the system catch this automatically next time?"
Two time rules. Per feature, plan plus review should take about 80% of the time and work plus compound about 20%. Across a developer's whole job, split 50/50 between shipping features and improving the system (review agents, documented patterns, test generators). Klaassen's arithmetic: one hour building a review agent saves about 10 hours of review over a year.
The plugin (house product)
The captured guide lists 26 agents, 23 commands and 13 skills. Key commands: /workflows:plan (parallel researchers plus a spec-flow analyzer), /workflows:review (14+ parallel reviewers such as security-sentinel, performance-oracle, data-integrity-guardian), /triage (human approve/skip per finding), /workflows:compound (six subagents write a searchable solution doc into docs/solutions/), and /lfg (the whole pipeline, 50+ agents, pausing only for plan approval). Note: [[2026-05-29-every-compound-engineering-upgrade]] covers a later version that grew the loop to more stages.
Beliefs to drop, beliefs to adopt
Drop eight beliefs: code must be hand-written; every line must be manually reviewed; solutions must come from the engineer; code is the primary artifact; writing code is the job; first attempts should be good (Klaassen puts first attempts at 95% garbage, second at 50%); code is self-expression; more typing means more learning.
Adopt: extract your taste into CLAUDE.md, agents, skills and commands; build safety nets instead of manual review ("if you don't trust the results, fix the system"); make the environment agent-native (the agent can run tests, read logs, take screenshots, open PRs); parallelize, because the bottleneck is now compute, not attention; and treat the plan as the new primary artifact.
Five-stage adoption ladder
0 manual; 1 chat-based copy-paste; 2 agentic tools with line-by-line approval (where most developers plateau); 3 plan-first, PR-only review (where compounding starts); 4 idea to PR on one machine; 5 parallel cloud execution with proactive agents. Each transition has a "compounding move": keep a prompt log (0→1), start CLAUDE.md (1→2), document what each plan missed (2→3), build a library of outcome-style instructions (3→4), document which work parallelizes and which is inherently serial (4→5). Skipping stages fails because trust has not been built.
Practical extras
- Three questions before approving any output: hardest decision made, alternatives rejected and why, what it is least confident about.
- Agent-native levels: basic dev (files, tests, commits) → full local (browser, logs, PRs) → production visibility (read-only logs, error tracking) → full integration (tickets, deploy).
- Skip permissions only on a branch with tests and easy rollback, never in production; git, tests, worktrees and PR review are the replacement safety net. Claimed gain: five to 10 times faster iteration.
- Team norms: silence is not plan approval; the initiator owns the PR whatever wrote it; human reviewers check intent and business logic, not syntax or security the agents already covered; handoffs state status, what is left and how to continue; "Ask Sarah" becomes "Sarah ran /compound."
- Beyond code: throwaway "baby app" design prototypes, design-taste and copy-voice skills, structured persona files, and release notes generated from the plan.
Mapping against Ray Data Co
The most concrete connection is the copilot agent factory: its 50+ stateless skills, document-tracked state, eval plans and human review gates are this loop, minus an explicit Compound step with a "would the system catch this next time?" check. Ray has the parts ([[2026-04-04-nightly-learn-and-ship-loop]], /improve, /self-review, memory feedback files) but the capture is ad hoc rather than a required close-out of every run.
- Review bandwidth is the binding constraint. The guide's answer matches RDCO's direction: replace human line review with parallel specialized critics plus P1/P2/P3 triage, and keep the human on intent. RDCO's fresh-eyes critics (verify-vault-write, station-critic, behavior-critic) are this, but findings are not yet ranked so that only P1s reach Ben.
- Ladder placement. On Klaassen's five stages Ray sits at 4 to 5 (always-on, multi-agent, crons, proactive). That fits Ben's Level 5 (Workflows) diagnosis on [[2026-06-02-every-eight-levels-ai-adoption]]: capability is high, and the remaining work is reliability scaffolding.
- Where RDCO deliberately departs. The skip-permissions advice assumes a sandbox. Ray runs on real accounts (iMessage, email, deploys), so RDCO keeps hard human gates on irreversible actions. The guide's own "never in production" caveat supports that line.
- "Silence is not approval" maps cleanly onto the decision-vs-action rule and the paper-trade deploy authorization standard.
Why it matters for RDCO / The Denominator
- Do: make "compound" a required last step in factory and Ray skill runs: write one tagged learning file per run and answer the "would we catch it next time" question in it. Rank critic findings P1/P2/P3 and only page Ben on P1.
- Do: run the three questions as a standard appendix on subagent reports. It is a cheap way to target his limited review time.
- Write: a Denominator piece that tests the 80/20 and 50/50 rules against enterprise agents. Hypothesis: successful enterprise agent teams spend about half their effort on the system around the agent (evals, review agents, solution docs), not on the agent. The adoption ladder plus "compounding moves" is a usable client diagnostic, but cite it as Every's framework.
⚠️ Sponsorship
The guide is house promotion for Every's free open-source Compound Engineering plugin, and it promotes Cora and other Every products. Every also sells AI consulting and runs Compound Engineering camps. Bias implications: claimed results (one-person teams, five to 10 times faster, 95% first-attempt garbage) are self-reported and unaudited, and the guide describes the plugin's own command set as the way to do the practice. The method stands without the plugin.
Related
- [[2026-04-04-compound-engineering]] - earlier vault summary of the methodology and plugin.
- [[2026-01-30-every-compound-engineering-framework]] - Shipper and Klaassen's framework essay.
- [[2026-05-29-every-compound-engineering-upgrade]] - later plugin version with an expanded loop.
- [[2026-07-25-every-opus5-compound-engineering-breakage]] - where the plugin's hand-offs broke under Opus 5.
- [[2026-04-04-nightly-learn-and-ship-loop]] - wrapping the Compound step in unattended automation.
- [[2026-01-28-every-stop-coding-start-planning]] - the planning half of the loop in depth.
- [[2026-01-26-every-claude-code-shipping]] - Klaassen's earlier parallel Claude Code workflow.
- [[2026-06-02-every-eight-levels-ai-adoption]] - the companion adoption ladder.
- [[feedback_verification_independent_worker_pattern]]