06-reference

every burn more tokens

2026-09-17·reference·source: Every·by Laura Entis
token-economicsagent-orchestrationmodel-selectionai-ops

Why this is in the vault

Every's internal playbook for deciding when expensive agentic token spend is worth it — and a concrete failure case (4.5B tokens, an "unruly swarm" of subagents) showing what happens without that discipline.

The core argument

Every CEO Dan Shipper doesn't cap experimentation spend; instead the team asks post-hoc: was the result worth what we spent, and could we reproduce it cheaper next time? Head of ops Arielle Shipper runs a lightweight early-warning system — a $500 auto-refill card triggering Slack pings on rapid charges — then checks the usage leaderboard and messages whoever's spiking. Her three standing questions: what did the run cost, what did it buy, what did we learn. The company is building a skill for personal AI spending P&Ls and per-employee model benchmarks (head of evals Mike Taylor) to make model/cost tradeoffs visible without discouraging ambitious bets.

The cautionary data point: head of video Randy Counsman burned 4.5 billion OpenAI tokens over a day and a half trying to get Astra to build a 3D face model via an orchestrator → implementer → subagents → judge-agent pipeline with an open-ended goal ("done extremely well"). An audit found the layered agents were passing growing context back and forth even for trivial status checks, burning tokens on coordination rather than work. He's since dropped the implementer layer, capped subagents at five, added checkpoints for his own feedback, and replaced the open-ended goal with a concrete target image for the judge agent to check against — a leaner setup with a reusable benchmark.

Secondary "steal this workflow" item: Spiral GM Marcus Moretti's model-routing discipline — route by task complexity (Sonnet for basic, Opus for scoped work, Fable 5.1 to plan larger features), then let the top model delegate execution to cheaper models and check their work, rather than running everything at max intelligence.

Mapping against Ray Data Co

Directly validates the founder's existing dispatch-hygiene rules rather than introducing a new idea: feedback_delegation_model_effort_pairing (pair model with effort, high/xhigh only for meaningfully large tasks) and feedback_verify_dispatch/verify-dispatch skill (checking for over-broad scope and missing acceptance contracts before sending a sub-agent dispatch) are exactly the guardrails that would have stopped Randy's failure mode — an orchestrator told to deploy subagents "as needed" with no acceptance contract and an ambiguous goal ("done extremely well"). The Every fix (cap subagent count, add a judge checkpoint against a concrete target instead of an open-ended instruction, drop the implementer layer) is a reusable pattern for any RDCO brigade-style build (station-spec-author → station-test-author → station-code-author → station-critic) that risks unbounded coordination overhead. Also reinforces feedback_no_batched_result_declaration in spirit — verify before scaling, don't assume high usage was productive.

Related

⚠️ Sponsorship

This is Every's own house content promoting its bundled products (Sparkle, Cora, Spiral, Monologue) and its Every All Access / Builder Pack subscription in the footer — self-promo, not a paid third-party sponsor. No bias implication for the substantive editorial content above; the footer promo is standard Every house-ad boilerplate.