"Agent-native Architectures" — Dan Shipper and Claude (Every guide)
Page meta: published 2026-01-17, last modified 2026-04-07.
Why this is in the vault
This is Every's canonical technical reference for building apps where the agent is a first-class user: five principles, a tool-granularity rule, a files-first state model, an approval matrix keyed to stakes and reversibility, and a named anti-pattern list. It is the design vocabulary behind the copilot agent factory and the Ray system, written down by someone else, so it is worth keeping at full length.
The core argument
Premise: a strong coding agent (an LLM with bash and file tools, looping until an objective is met) turns out to be a strong general-purpose agent. So features stop being code you write and become "outcomes you describe," reached by an agent with tools in a loop. Sections that are Claude's untested contributions carry a "Needs validation" callout (naming conventions, files vs. database, the approval matrix, iOS checkpointing, capability discovery).
The five principles (each with a test)
- Parity - whatever a user can do in the UI, the agent can achieve through tools. Test: pick any UI action; can the agent reach that outcome? A capability map (user action -> tool path) is the audit artifact.
- Granularity - tools are atomic primitives; features are prompts. Test: to change behavior, do you edit a prompt or refactor code? A
classify_and_organize_filestool has your judgment baked in;read_file/move_fileplus a prompt leaves judgment with the agent. - Composability - with parity and atomic tools, a new feature (e.g. a "weekly review") is just a new prompt.
- Emergent capability - users ask for things you never built; the agent composes tools or fails, and both outcomes are signal. A six-step flywheel: build primitives, observe unanticipated asks, watch success or failure, spot patterns, add domain tools or prompts, repeat.
- Improvement over time - apps improve without shipping code, through accumulated context files, developer prompt updates, and user prompt edits. Self-modification is "emerging" and needs approval gates, checkpoints, rollback, and health checks.
Primitives to domain tools to code
Start with pure primitives. Add domain tools deliberately for three reasons only: vocabulary (a create_note tool teaches what a note is), guardrails (validation that should not rest on judgment), and efficiency. Rule: a domain tool is one conceptual action from the user's view; it may validate mechanically, but the decision of whether to act belongs in the prompt. "Domain tools are shortcuts, not gates" - keep primitives reachable unless security or integrity says otherwise. Hot paths can graduate to deterministic code, but the agent must still be able to trigger them and fall back to primitives.
Files as the universal interface
Files are already fluent to agents, inspectable, portable, and self-documenting. Heuristic: if a human can read your folder structure, an agent can too. Suggested: entity-scoped directories, markdown for human content, JSON for structure, plus a per-agent context.md (who I am, what I know about the user, what exists, recent activity, guidelines, current state) read at session start and updated as state changes. Files for legibility, databases for high-volume relational data; a hybrid keeps a file "source of truth" synced to a DB. Where agent and user write the same files you need a conflict model: last-write-wins, check-before-write, separate spaces (agent writes to drafts/, user promotes), append-only logs, or locking.
Execution patterns
- Explicit completion signals - a tool result carries success/failure and a separate continue/stop flag; never infer "done" from heuristics. Pause, escalate, and retry signals are flagged as not yet standard.
- Model tier selection - choose fast/balanced/powerful per agent deliberately; do not default to the biggest.
- Partial completion - task-level status (pending, in progress, completed, failed, skipped) with notes, checkpointed so a resume continues mid-list.
- Bounded context - summary-to-detail-to-full tools, mid-session consolidation, assume the window fills.
- Shared workspace and context injection - agent and user share one data space; the system prompt lists available resources, capabilities, and recent activity. "No silent actions": stream thinking, tool calls, and task progress, because "silent agents feel broken."
Product implications
Progressive disclosure (Excel and Claude Code: simple entry, no ceiling) and latent demand discovery - the agent becomes a research instrument; failed requests reveal tool or parity gaps. The approval matrix: low stakes/easy reversal auto-applies; low/hard gets a quick confirm; high/easy is suggest-and-apply; high/hard (sending email) needs explicit approval. An explicit user request already counts as approval. Self-modification must be legible: visible, understood, reversible.
Mobile sections (iOS-specific) cover checkpoint/resume after every tool result and a server orchestrator for long jobs. Advanced: dynamic capability discovery (list_available_types + read_data(type) instead of 50 endpoint tools) and a CRUD audit per entity.
Named anti-patterns
Agent as router; build the app then bolt on an agent; request/response thinking; defensive tool design (strict enums everywhere); happy path in code; workflow-shaped tools; orphan UI actions; context starvation; gates without reason; artificial capability limits; static mapping where discovery fits; heuristic completion detection. The closing test: describe an in-domain outcome you never built a feature for. If the agent loops to success, it is agent-native.
Mapping against Ray Data Co
The most concrete hit is the copilot agent factory: 50+ stateless skills over document-tracked state is the guide's "atomic tools plus files as state plus features as prompts" architecture almost line for line, which gives Ben an external citation for a design he arrived at independently.
- Gaps it exposes in the factory. Three checks the factory should run: a CRUD completeness audit per tracked entity (can the agent update and delete, not only create?); explicit completion signals rather than "output file exists" heuristics, which is exactly the fragile pattern the guide names; and a per-skill model tier choice written down rather than defaulted.
- Ray already matches the approval matrix. RDCO's gates (reversible vault and Notion writes auto-apply, deploys and external email are human-gated, founder sends) are the stakes-by-reversibility grid.
~/.claude/state/working-context.mdis thecontext.mdpattern. The vault-as-files choice is the files-first argument. - Where RDCO diverges on purpose. The guide pushes "default to open" and warns against gates without reason. RDCO deliberately runs more gates because the binding constraint is review bandwidth, not capability. That is a defensible difference to name, not a flaw: the guide is written for single-user consumer apps where the user is the reviewer.
- Bias caveat. About half of the operational detail is labeled as Claude's unvalidated suggestion, and the mobile half is iOS-specific to Every's own apps.
Why it matters for RDCO / The Denominator
- DO: add a one-page "agent-native audit" to the factory's eval plan: parity map, CRUD per entity, explicit completion tool, model tier per skill, context injection check. Each is a yes/no a client reviewer can verify.
- DO: adopt the "separate spaces" conflict model formally for Ray (agent drafts, human promotes) and write it into the review-gate SOP; it also cuts review load because drafts are scoped.
- WRITE: a Denominator piece on where enterprise agents should break this guide's "default open" rule. Thesis: in enterprises, gates are justified by review bandwidth and blast radius, and successful deployments make gating a recorded decision rather than a reflex. The guide supplies the vocabulary; Ben supplies the production counterexamples.
⚠️ Sponsorship
House promo, no paid third party. The page offers "Use in compound engineering" and links Every's own apps (Reader, Anecdote) as evidence. The bias is mild: the principles stand alone, but the evidence base is Every's own small, single-user apps, and the coauthor is the model whose SDK the guide recommends.
Related
- [[2026-02-17-every-build-agent-native]] - Katie Parrott's companion piece: lessons from Cora, Spiral, Sparkle, Monologue; "start with three tools."
- [[2026-06-16-every-agent-native-tool]] - Every applying these principles to an internal tool built by non-engineers.
- [[2026-05-05-every-codex-native-apps]] - same agent-native frame applied to Codex-native apps.
- [[2026-02-21-koylanai-personal-brain-os]] - independent arrival at "the file system is the new database."
- [[2026-07-04-addy-osmani-loop-engineering]] - loop, skills, connectors: the same outcome-in-a-loop shape from Google's side.
- [[2026-02-09-every-compound-engineering-guide]] - the plugin this guide points readers to.
- [[2026-06-02-every-eight-levels-ai-adoption]] - Level 5 "Workflows" is where these reliability patterns live.
- [[2026-05-03-every-agent-native-product-management-guide]] - the PM layer on top of an agent-native build.