Codex Graded My AI Habits. Then It Became My Coach.
Why this is in the vault
A concrete, five-step personal protocol for using an AI agent to diagnose and systematically level up your own usage of that same agent, including a recurring weekly grading cadence and a "diagnose the failure, propose the smallest durable fix, review before saving" correction loop.
The core argument
Arielle Shipper, Every's head of operations, used to run projects as single long chat threads. She asked Codex to grade her against Every's own "eight levels of AI adoption" framework by reading the framework article and scanning her actual session history — Codex placed her at 5.5 (solid agent use, but one agent per job rather than deputizing specialists). She then asked Codex what would move her to a 6 or 7; it pointed to a real finished project (cross-checking Stripe transaction data for sales-tax purposes) and showed how it could have been split across a schema mapper, API extractor, workbook builder, and tax-logic verifier. When the jargon ("orchestrator," "context packet") didn't mean anything to her, she asked Codex to explain the mechanics in plain terms and to walk her through redoing a past project at the next level. She then made the grading recurring — a standing weekly task that scans the week's sessions, scores her, and lists specific experiments for the following week — which took her from 5.5 to 7.9 over four months. Finally, when a multi-agent finance task produced correct output (Ramp card limits) but wrong tone (Slack messages that read like corporate memos despite a saved communication-style skill), she had Codex diagnose why the correction happened, articulate what "good" looks like, and propose the smallest durable instruction change — reviewed by her before being saved — via a /self improve-style loop.
Mapping against Ray Data Co
Concrete connection: Shipper's /self improve loop — the agent reads the original output plus the human's corrections, diagnoses the failure mode, proposes the smallest durable instruction change, and the human reviews before it's encoded — is structurally the same mechanism as RDCO's own /improve skill and the feedback_implementation_notes_sub_agent_pattern discipline (implementing sub-agents keep a running notes file so fixes persist past a single session). The piece is a useful external validation that this "capture the correction, don't just fix and move on" pattern is a generalizable practice, not an RDCO-specific eccentricity.
The sharper, less-covered idea is the recurring self-grading cadence — a standing weekly task that scans the agent's own session history and scores the human's delegation habits against a fixed rubric, with specific next-week experiments attached. RDCO has the inverse of this today: fresh-eyes critics (feedback_verification_independent_worker_pattern, feedback_workflow_agent_output_integrity) that grade artifacts before the founder sees them, but nothing that grades Ray's own delegation habits over time the way Shipper's weekly report grades hers — e.g., whether dispatches are over-scoped, whether feedback_delegation_model_effort_pairing is actually being followed, whether sub-agent context packets are as minimal as the orchestrator/specialist model calls for. That's a gap worth naming rather than a technique to copy outright, since Ray's harness already has /self-review and /vault-health running related but artifact-scoped checks.
⚠️ Sponsorship
No named third-party sponsor. The standard Every footer runs house self-promotion: an "Every All Access" / Builder Pack upsell and a bundled-software pitch for Every's own products (Sparkle, Cora, Spiral, Monologue). Flagged per RDCO convention (sponsor_entity: house-promo) as a commercial incentive embedded in editorial content, not paid third-party advertising.
Related
- [[2026-09-20-every-copy-our-homework]]
- [[2026-09-25-every-copilot-org-chart-autopilot]]
- [[feedback_implementation_notes_sub_agent_pattern]]
- [[feedback_workflow_agent_output_integrity]]
- [[feedback_delegation_model_effort_pairing]]
- [[project_channels_agent_setup]]