06-reference

every to read or not to read the code

2026-09-08·reference·source: Every·by Kieran Klaassen
skill-erosionagentic-codingdiscernmentverificationcompound-engineering

"To Read—Or Not to Read the Code?" — Every (Kieran Klaassen)

Why this is in the vault

Klaassen (GM of Cora, Every's email product) names the exact tension RDCO lives inside every day: delegating mechanical work to agents is correct, but if the human stops reading code entirely, the human's judgment erodes even as the system's capability compounds — and it's the human's judgment that catches the failure mode an agent can't see coming.

The core argument

Klaassen has spent a year hunting for parts of coding that still need him and replacing himself with agents — the essence of "Compound Engineering." It worked: he ships more than ever. But he noticed his own mind going soft, "like a TikTok feed" of short dopamine hits that leave nothing behind. So he built a second loop, separate from shipping: reading code again, not to verify it (planning/testing/review agents do that better than he does) but to learn.

His live example: a customer reported Cora permanently deleted sent emails. He'd built /ce-explain — a command that traces a concept through his own codebase — and had already used it once to build a mental model of how a Gmail change flows through Cora. That model let him tell a plausible-sounding fix from a correct one. The bug: Cora drafted a message, saved its ID to clean up later, but the customer edited and sent that same draft through Gmail directly — so Cora's cleanup deleted a message that was no longer a draft. The first two proposed fixes (wait 30 seconds; skip deletion if Cora sent it) were both guesses that didn't match what actually happened. The fix that held: before deleting anything, ask Gmail directly whether it's still a draft, and if Gmail can't say for certain, delete nothing — a rule instead of a guess, validated when the same failure mode recurred hours later and the guard held.

He cites a paper by Margaret Mitchell, Avijit Ghosh, and Samir Passi (arXiv 2608.23642) finding that extended AI-agent use measurably erodes the vigilance, critical thinking, and domain skill that human oversight depends on — the "irony of automation," a term from 1983 safety research on automated factories and power plants. Agents just made the erosion faster.

His four countermeasures: (1) revisit merged PRs you didn't read line-by-line and ask the model why it chose an unfamiliar pattern — keep a list of gaps as a syllabus; (2) ask for the mechanics (where does the process start, what touches it) rather than just the diff; (3) ask why a design choice exists — the code shows what the system does, not why someone built it that way, and you can't tell if removing a safeguard is simplification or a regression waiting to happen; (4) let the model quiz you after a session (borrowed from Thariq Shihipar) without turning the quiz into a merge gate — tests gate the merge, the quiz finds what you still need to learn.

His answer to "I only care that it works": on an AI product, a taste question is often a mechanics question in disguise — deciding how much context an LLM call gets, how long a customer waits, what a request costs, is a technical tradeoff, and you can't evaluate a proposed limit as real vs. arbitrary without understanding it.

Mapping against Ray Data Co

Direct hit against the founder's own PR-only workflow discipline (feedback_pr_only_workflow): branch-and-PR keeps Ray from writing to main directly, but says nothing about whether the founder is reading those PRs to keep his own judgment sharp, versus trusting the autonomous review/merge gate entirely. Klaassen's distinction — verification is an agent's job, learning is the human's separate job that has to be forced — maps cleanly onto why RDCO runs fresh-eyes critics (verify-dispatch, verify-strategic-output, verify-vault-write) for correctness and still needs the founder occasionally reading a diff or a brigade-station output for reasons that have nothing to do with catching bugs. His Gmail-draft bug — where the first two plausible-sounding fixes were both wrong because nobody understood the actual mechanics — is a direct analogue to the "verify blockers against the source, not notes" and "static asset diff" feedback memories: a guess that sounds right (30-second wait) is exactly the failure mode that a hard-checked rule (ask the source system directly) prevents. It also sharpens the founder's "advisor not pair programmer" posture (feedback_advisor_not_pair_programmer): Klaassen's point cuts the other way from that memory — an advisor role that never reads the artifact is at risk of the same discernment decay he's describing, so the boundary should be judgment-on-demand, not zero-exposure.

Related

[[2026-04-20-every-ai-autopilot-verification-decay]] [[2026-08-03-data-engineering-central-keeping-technical-skills-llm-age]] [[2026-06-03-commoncog-software-dark-factory-qa-reflections]] [[feedback_pr_only_workflow]] [[feedback_advisor_not_pair_programmer]]

⚠️ Sponsorship

No paid third-party sponsor block in this issue. The only promotional content is the standard Every house footer: Every All Access / Builder Pack upsell, and cross-promo for Every's own product bundle (Sparkle, Cora, Spiral, Monologue). Klaassen himself is the GM of Cora, one of the cross-promoted products, and the essay's examples are drawn entirely from Cora's own incident history — a self-promo angle (marked sponsor_entity: self) rather than a paid third-party placement. Consistent with the README's standing note that Every's one confirmed paid sponsor (Brief/briefhq.ai) remains a one-off, not a recurring pattern.