06-reference

every agent native architectures guide

2026-01-17·reference·source: Every·by Dan Shipper (coauthored with Claude)
every-greatest-hitsagent-nativeagent-architecturetool-designfiles-as-interfaceapproval-gatescontext-engineering

"Agent-native Architectures" — Dan Shipper and Claude (Every guide)

Page meta: published 2026-01-17, last modified 2026-04-07.

Why this is in the vault

This is Every's canonical technical reference for building apps where the agent is a first-class user: five principles, a tool-granularity rule, a files-first state model, an approval matrix keyed to stakes and reversibility, and a named anti-pattern list. It is the design vocabulary behind the copilot agent factory and the Ray system, written down by someone else, so it is worth keeping at full length.

The core argument

Premise: a strong coding agent (an LLM with bash and file tools, looping until an objective is met) turns out to be a strong general-purpose agent. So features stop being code you write and become "outcomes you describe," reached by an agent with tools in a loop. Sections that are Claude's untested contributions carry a "Needs validation" callout (naming conventions, files vs. database, the approval matrix, iOS checkpointing, capability discovery).

The five principles (each with a test)

  1. Parity - whatever a user can do in the UI, the agent can achieve through tools. Test: pick any UI action; can the agent reach that outcome? A capability map (user action -> tool path) is the audit artifact.
  2. Granularity - tools are atomic primitives; features are prompts. Test: to change behavior, do you edit a prompt or refactor code? A classify_and_organize_files tool has your judgment baked in; read_file/move_file plus a prompt leaves judgment with the agent.
  3. Composability - with parity and atomic tools, a new feature (e.g. a "weekly review") is just a new prompt.
  4. Emergent capability - users ask for things you never built; the agent composes tools or fails, and both outcomes are signal. A six-step flywheel: build primitives, observe unanticipated asks, watch success or failure, spot patterns, add domain tools or prompts, repeat.
  5. Improvement over time - apps improve without shipping code, through accumulated context files, developer prompt updates, and user prompt edits. Self-modification is "emerging" and needs approval gates, checkpoints, rollback, and health checks.

Primitives to domain tools to code

Start with pure primitives. Add domain tools deliberately for three reasons only: vocabulary (a create_note tool teaches what a note is), guardrails (validation that should not rest on judgment), and efficiency. Rule: a domain tool is one conceptual action from the user's view; it may validate mechanically, but the decision of whether to act belongs in the prompt. "Domain tools are shortcuts, not gates" - keep primitives reachable unless security or integrity says otherwise. Hot paths can graduate to deterministic code, but the agent must still be able to trigger them and fall back to primitives.

Files as the universal interface

Files are already fluent to agents, inspectable, portable, and self-documenting. Heuristic: if a human can read your folder structure, an agent can too. Suggested: entity-scoped directories, markdown for human content, JSON for structure, plus a per-agent context.md (who I am, what I know about the user, what exists, recent activity, guidelines, current state) read at session start and updated as state changes. Files for legibility, databases for high-volume relational data; a hybrid keeps a file "source of truth" synced to a DB. Where agent and user write the same files you need a conflict model: last-write-wins, check-before-write, separate spaces (agent writes to drafts/, user promotes), append-only logs, or locking.

Execution patterns

Product implications

Progressive disclosure (Excel and Claude Code: simple entry, no ceiling) and latent demand discovery - the agent becomes a research instrument; failed requests reveal tool or parity gaps. The approval matrix: low stakes/easy reversal auto-applies; low/hard gets a quick confirm; high/easy is suggest-and-apply; high/hard (sending email) needs explicit approval. An explicit user request already counts as approval. Self-modification must be legible: visible, understood, reversible.

Mobile sections (iOS-specific) cover checkpoint/resume after every tool result and a server orchestrator for long jobs. Advanced: dynamic capability discovery (list_available_types + read_data(type) instead of 50 endpoint tools) and a CRUD audit per entity.

Named anti-patterns

Agent as router; build the app then bolt on an agent; request/response thinking; defensive tool design (strict enums everywhere); happy path in code; workflow-shaped tools; orphan UI actions; context starvation; gates without reason; artificial capability limits; static mapping where discovery fits; heuristic completion detection. The closing test: describe an in-domain outcome you never built a feature for. If the agent loops to success, it is agent-native.

Mapping against Ray Data Co

The most concrete hit is the copilot agent factory: 50+ stateless skills over document-tracked state is the guide's "atomic tools plus files as state plus features as prompts" architecture almost line for line, which gives Ben an external citation for a design he arrived at independently.

Why it matters for RDCO / The Denominator

⚠️ Sponsorship

House promo, no paid third party. The page offers "Use in compound engineering" and links Every's own apps (Reader, Anecdote) as evidence. The bias is mild: the principles stand alone, but the evidence base is Every's own small, single-user apps, and the coauthor is the model whose SDK the guide recommends.

Related