"Where Do You Fall on the Eight Levels of AI Adoption?" - Mike Taylor (guide version with Laura Entis)
Re-read at full length 2026-10-03 (paid).
Why this is in the vault
A clean eight-rung maturity ladder for AI adoption (chatbot through manager-of-sub-agents) that gives RDCO a shared vocabulary to locate where Ben and Ray-the-COO-agent sit and what the next rung actually requires. The earlier version of this note (~740 words) was built from the /p/ essay and only named the rungs; the full guide adds per-level sample prompts, the human-judgment step at each rung, and explicit "when to move up" signals, folded in below.
The core argument
Taylor's correction up front: chasing the viral power user ("12 Claude Code sessions in parallel") is the wrong target. He built the ladder because Steve Yegge's "Gas Town" framing didn't match his client experience. The right level for a task is set by how much you trust the AI to work unsupervised and how costly a mistake is. For high-stakes work you either stay low and supervise, or invest the engineering, time, and tokens to get the same quality with less oversight. Most people struggling to adopt have good reasons: output quality is too low, or too expensive to raise. Rough placement today: knowledge workers sit at Levels 1-4; engineers at 5-8, because they can build scaffolding before platforms ship it.
The ladder, with what the guide adds per rung
| Level | What it is (tools named) | Human judgment | Signal to move up |
|---|---|---|---|
| 1 Chatbot | You ask, it answers (ChatGPT, Claude, Gemini) | Confirm tone and facts | You're tired of copy-pasting context in and results out |
| 2 Copilot | AI inside your file (Cursor, Claude in Excel) | Does it sound like you; are formulas right | You need to work across multiple sources, not one file |
| 3 Agent | Executes step by step with approvals; reactive (Cowork, Codex) | Approve plan, fix interpretation | You want to trade control for speed or one-shot a prototype |
| 4 Autopilot | Skip approvals, review the end result (Lovable, Claude Code) | Test it; note where reliability is needed | Results are fast but uneven and the work is higher-stakes |
| 5 Workflows | A harness that plans, self-reviews, and runs safeguards (compound engineering, Claude Workflows) | Is the confidence score justified | Activating the workflow yourself is now the bottleneck |
| 6 Assistant | Proactive, unprompted, background (OpenClaw, Hermes, Claude Managed Agents) | Decide what is urgent; refine rules | You want more, but one agent's memory is overloaded |
| 7 Multi-agent | Several long-running agents, each with a role | Review each agent's output and scope | You lose track of which agent owns what |
| 8 Orchestrator | A manager agent runs sub-agents (Gas Town, Paperclip, Symphony) | Set the escalation threshold | (top rung; "highly experimental") |
Sample prompts worth stealing. L1: draft a client follow-up but "tell me if anything sounds unclear or unsupported before you start." L3: update a board deck from a spreadsheet, proposing edits slide by slide and flagging inconsistent source data. L4: a demo-ready lead-scoring tool on dummy data. L5: run /ce-plan with files, edge cases, and verification before code, then a review prompt asking for a 1-100 confidence score, the weakest parts, and another pass "until you are above 90" or an explanation why not. L6: every 30 minutes, flag meetings in the next two hours that need prep and draft missing agendas. L7: a second always-on agent (the guide's example is literally one running on its own Mac mini) gets a separate job, kept apart so contexts don't bleed. L8: a pipeline that reviews submissions against standards and runs tests, escalating "only when they require a judgment call."
Other new detail. Level 5 discipline will be "baked into platforms over the next six to 12 months." Every's own consulting team runs an L6 assistant for project management and sales pipeline that works only because a senior engineer maintains it. OpenClaw is called "inherently unstable," with unsolved memory. Risk tolerance matters more at L6 than any earlier rung. Closing analogy: you wouldn't brag about eight interns working overnight whose output you never checked; expect months of supervised work before trusting an agent at the next level.
Mapping against Ray Data Co
Strong, and unusually precise; this ladder is almost a self-assessment built for RDCO.
- Where Ray-the-COO already sits on capability: rungs 6-8. The always-on Mac mini agent works unprompted on crons (morning-prep, open-threads-check, sync-contacts, curiosity), which is Level 6. The sub-agent fan-out (one per article, per research question, per critic axis) is Level 8 in shape. RDCO is not climbing toward multi-agent; it already operates there on specific surfaces.
- The L4→L5 tension in his vocabulary. The founder places RDCO at L4-L5 (see [[2026-05-06-dec-ai-not-replacing-curious-developers]]); his current self-diagnosis is Level 5. The unhobbling work is Level 5 labor: systems that make output reliable (independent verification, fresh-eyes critics, the 12-stage production workflow). RDCO has high-rung capability while rung-5 reliability scaffolding is the real bottleneck. The next unlock is reliability, not more autonomy.
- Trust-and-blast-radius is already encoded as hard gates. Deploys and external email stay human-gated; reversible work (Notion adds, vault notes, paper-trade drafts) runs autonomously. The ladder is a clean external articulation of the IC-mode vs production-mode split.
- phData tie-in. The ladder, now with prompts and move-up signals, is a ready-made client diagnostic for the copilot agent factory: place each team's use case on a rung, then sequence them up. Each signal is an interview question.
Honest caveat: the ladder is a vocabulary, not a measurement. It flatters by giving RDCO high-rung labels; the unglamorous question is rung-5 reliability.
What Level 5 → 6 requires, given Ben's review bandwidth
The guide's L5 exit signal is "activating the workflow yourself is the bottleneck." Ben's actual bottleneck is different: output already outruns his review. Climbing to L6 by adding more unprompted work makes that worse, because every L6 run still lands in his review queue. Read through the guide, the real L6 prerequisite is the human-judgment column shrinking: the L8 pipeline example works because review escalates only judgment calls. So 5 → 6 for Ben means:
- Calibrated self-review before anything reaches him. Make the L5 confidence-loop prompt a standard critic step, and track whether the critic's scores predict his edits. A score he can't trust does not save review time.
- Tier by blast radius, review by exception. Reversible, low-stakes outputs (vault notes, digests) ship with sampling audits instead of per-item review; only irreversible or public outputs queue for him.
- One escalation queue, not many outputs. L6 assistants should deliver a short "needs your call" list, the guide's model, rather than finished artifacts.
- Measure review minutes per shipped artifact and promote a workflow to L6 only once that number falls while his edit rate stays flat.
Until those hold, staying at L5 on high-stakes work is the guide's own advice, not a failure.
Why it matters for RDCO / The Denominator
- Diagnose with the signals, not the labels. For phData clients, turn the "when to move up" column into a short intake questionnaire.
- The Denominator angle: "successful enterprise agents" mostly live at L4-L5 with tight escalation design, not at L8. A post arguing that the review queue (the denominator of agent output) decides the ceiling would be distinctive against Every's ladder.
- For Ray: instrument review minutes and critic-score accuracy before adding any new cron.
⚠️ Sponsorship
House promo, no third-party sponsor. Taylor is Every's head of tech consulting, and the guide closes with a pitch for Every's consulting team; the /p/ post footer also plugs Spiral, Sparkle, Cora, Monologue, and Proof. The L5 prompts route through Every's compound engineering commands (/ce-plan, /ce-code-review). Bias implication: the framework doubles as a top-of-funnel diagnostic for Every consulting and nudges Level 5 toward Every's own tooling; the rungs and signals stand on their own.
Related
- [[2026-05-21-every-after-automation]] - Dan Shipper on what comes after the automation rungs; same Every house view on orchestration.
- [[2026-05-06-dec-ai-not-replacing-curious-developers]] - the L4→L5 unhobbling thesis this ladder reframes (developer-as-supervisor).
- [[2026-05-15-every-team-agents-vs-personal-pets]] - Every's own employees climbing these rungs in practice.
- [[2026-05-26-every-codex-for-knowledge-work-power-user-guide]] - adjacent Every guide on pushing knowledge work up the ladder.
- [[2026-01-20-every-ai-teaching-management]] - prior Mike Taylor piece; managing agents as a management skill.
- [[2026-02-09-every-compound-engineering-guide]] - the Level 5 harness the guide points to.
- [[2026-03-03-every-openclaw-comprehensive-guide]] - the Level 6 platform it calls unstable.
- [[2026-04-20-every-ai-autopilot-verification-decay]] - why review-by-exception needs calibrated checks.
- [[2026-06-07-alphasignal-async-agents-release-bottleneck]] - human review as the ceiling on agent throughput.