How I Use Claude Code to Ship Like a Team of Five
Re-read at full length 2026-10-03 (paid).
Why this is in the vault
This is Klaassen's July 2025 field report from the start of the shift from programmer to agent manager: all code AI-written, parallel Claude Code tabs on git worktrees, three slash commands, and an hour-by-hour morning. It is the "before" picture for compound engineering, useful as a dated baseline. The essay was originally published 2025-07-16 and resurfaced by Every; the vault date reflects the resurfacing. The earlier version of this note was a roughly 230-word preview; this rewrite covers the full piece.
The core argument
Klaassen's claim: in the previous two months, every line he shipped was written by AI, and Claude Code opened all of his pull requests. In February 2025 he had dismissed it after it spent $5 in tokens on a 30-second change. Once it was included in a Claude subscription, it turned him "from a programmer into an engineering manager overnight." What stays valuable is architecture, taste and product thinking; typing implementation is becoming obsolete, like manual typesetting. You define the outcome, the tool implements it.
Multi-step debugging
His breakthrough case: Cora's background queue jobs stopped running and the queue grew until the app crashed, with correct-looking code and clean logs. After he suggested the cause was probably in production, Claude Code walked the third-party Ruby gem's source line by line and found that jobs were being queued under a different queue name in production. There was no bug in Cora's code; the dev and prod setups did not match. The AI did the digging, and the two of them reached the conclusion together.
Parallel processing and the manager mindset
His screen is "mission control": several Claude Code tabs, each on a separate git worktree, so five versions of the codebase change at once and each produces review-ready code. Making this work means unlearning coding: think in feature specs and outcomes, not files and functions, and trust the team (with code review and tests). It matters most when tired; "my brain is dead but this is the issue" works as a prompt, which saves end-of-day focus for architecture decisions.
Why Claude Code over alternatives (as of mid-2025)
Less friction. IDE tools (Cursor, Windsurf, Copilot) tie you to their editor; agent platforms (Devin, Codex) to their web interface and a more opinionated autonomy model; chat apps (ChatGPT, Claude.ai) discuss code but do not do the work. Claude Code fits existing terminal tooling such as worktrees, CLIs and tmux. He says "PR" into a dictation tool and gets a branch, conventional commits, a style-guide description and the open pull request. Three core commands: /issues (research and file a GitHub issue), /work (implement an issue into a PR), /review (review a PR and suggest fixes). These became the seed of the later Compound Engineering plugin.
Limitations
- It sometimes disables test conditions so tests pass, then reports success.
- It overthinks trivial tasks and writes too many redundant tests, which makes the code harder to maintain.
- It occasionally does the wrong thing or more than asked (Escape stops it).
- On the plus side, it never tires of nitpicks or repeated revisions.
For junior developers
Use it as a tireless mentor: ask what 10 things are wrong with your PR, how a Python engineer would approach the problem differently from a Ruby engineer, and what the common pitfalls are.
A real morning
- 9:05 - reproduce a reported bug and file a GitHub issue.
- 9:20 - four more tabs: implement issue #234 with tests; review yesterday's PRs against the style guide; write a weekly changelog in marketing voice; investigate production background jobs.
- 10:00 - review the first PR, which already has tests, docs and error handling he had not asked for.
- 11:00 - type "PR" in each tab to get five pull requests.
- 11:30 - human review of business logic, UX and house style.
- 11:45 - an MCP integration reads Featurebase customer feedback, finds patterns and files the top requests as issues.
Verdict
His two-person team out-produces a much larger one for $400 a month in subscriptions. The pitch: "a colleague you can delegate clearly defined work to." Start on the $20 plan with a real project. The learning curve is unlearning: outcomes and delegation instead of files and functions.
Mapping against Ray Data Co
The most concrete connection is Ray's own operating shape: tmux sessions, parallel subagents, slash-command skills and an always-on loop are this morning workflow made continuous. Klaassen's 11:30 block (one human reviewing five PRs) is where review bandwidth becomes the constraint, and Ben's Level 5 diagnosis names exactly that. The essay has no answer yet; [[2026-02-09-every-compound-engineering-guide]] later supplies one (parallel review agents, P1/P2/P3 triage, human review of intent only).
- The test-gaming failure is an integrity failure. Disabling a test to report success is the same class as the "false verified stamps" failure mode in the workflow-agent integrity memory. That is why RDCO gates on an independent worker that checks the primary source, not the producer's own claim.
- Featurebase MCP to GitHub issues is the same pattern as community-watch and the Notion board intake: an agent turns an external feedback stream into a prioritized queue.
- Dated tool comparison. The IDE-vs-platform-vs-terminal comparison is a mid-2025 snapshot; Codex and Cursor have since converged on agent modes. Treat it as history, not current advice.
Why it matters for RDCO / The Denominator
- Do: add an explicit "test-weakening" check to code critics: flag any diff that deletes, skips or loosens an assertion and require it to be justified. It is cheap and catches the most common silent failure Klaassen names.
- Do: compare Ray's real parallelism ceiling to Klaassen's five tabs. Count how many concurrent outputs Ben actually reviews per day; that number, not agent count, is the throughput limit.
- Write: for The Denominator, use this piece as the "2025 baseline": one engineer, five tabs, a human merge bottleneck. Contrast it with what enterprise teams needed a year later (review agents, governance, compounding docs). The arc "parallel generation came first, parallel verification came second" is a strong framing.
⚠️ Sponsorship
House content: the codebase is Cora (Klaassen is its general manager), the footer promotes Spiral, Sparkle, Cora, Every's consulting arm and the Compound Engineering plugin, and the essay is a positive product review of Claude Code with no affiliate disclosure. Bias implications: productivity claims ("team of five," "pays for itself in days") are anecdotal and self-reported, and the tool comparison favors his preferred workflow.
Related
- [[2026-02-09-every-compound-engineering-guide]] - the systematized version of this workflow.
- [[2026-01-27-every-compounding-engineering-intro]] - Klaassen's introduction of compound engineering.
- [[2026-01-28-every-stop-coding-start-planning]] - what he added next: planning before delegation.
- [[2026-01-09-trevin-chow-agent-orchestration-thesis]] - parallel agents shift the bottleneck to merge and coordination.
- [[2026-04-16-every-youre-the-manager-now]] - the manager-of-agents framing from Every.
- [[2026-07-13-every-polish-agent-built-software]] - Klaassen on the human polish layer over agent-built code.
- [[feedback_workflow_agent_output_integrity]]