Why this is in the vault
Concrete worked example of a compounding self-improve skill loop (KateBench) plus a multi-item signal roundup (open-weight model economics, AI-authorship debate) directly relevant to how RDCO builds and routes its own skills.
The core argument
Every CEO Dan Shipper's team built KateBench, a copyediting agent trained on editor Kate Lee's 30,000+ historical edits, only viable once GPT-5.6 shipped. Every article now runs through it before Kate's human review; her accept/reject/modify decisions, tracked in Google Docs, feed back in and Codex rewrites the skill so it compounds. The team is now testing whether this generalizes past copyediting — DanLens (cloning Dan's marketing instincts) is the first extension — but flags that UX and judgment work produces a messier training signal than copyedits do. Not every kind of expertise is equally "clone-able" yet.
Curation section
Discuss — AI authorship on op-eds
Stanley Druckenmiller confirmed using AI to draft a WSJ op-ed criticizing Treasury Secretary Bessent; WSJ opinion editor Paul Gigot defended publishing it as reflecting Druckenmiller's "genuine opinion," contrasting with the Financial Times' outright AI ban for columnists. Sparked debate over whether a byline certifies authorship or just ownership of an argument.
Skill Share — self-improving Codex skill
Every's Arielle Shipper built a skill that, after Codex makes a mistake, interrogates what went wrong and proposes a targeted edit to Codex's own operating instructions to reduce recurrence — open-sourced on GitHub.
Signal — an opening for open-weight models
Ramp spending data shows Fable captured only ~6% of business Anthropic-token spend a month post-launch, partly on enterprise data-retention/security concerns, partly because most knowledge work doesn't need frontier intelligence. Open-weight models' share of tokens routed through Vercel's AI Gateway rose from 11% (April) to 29% (June) — cheap-and-good-enough models are absorbing routine-task volume even as spend stays concentrated on frontier models.
AI & I podcast — Walleye Capital
The $10B hedge fund Walleye Capital makes AI fluency mandatory for all 400 employees; CEO Will England treats AI adoption like using the internet in 1995 — non-negotiable — and grades on results, not effort put in.
The Daily Driver
Every staff report their current model picks for the week: a mix of GPT-5.6 Sol, Opus 5, Grok 4.6, and Fable depending on task and reasoning-effort level — an informal weekly model-routing survey.
Mapping against Ray Data Co
KateBench is the same shape as RDCO's own /improve skill and the implementation-notes sub-agent pattern: capture a human's accept/reject decisions on agent output, feed them back to rewrite the skill's own instructions, and let it compound. Every's own open question — that UX/judgment work is too messy to KateBench-ify yet — is a live one for RDCO's DESIGN-
⚠️ Sponsorship
This issue carries a paid placement for Grok Bot (xAI), pitched as autonomous cloud-hosted "teammate" bots ("Scout," "Olivia," "Talent Scout") with a "Meet the Bots" CTA and newsletter-tracked UTM params — a new one-off sponsor, not the previously observed Brief/briefhq.ai. The article content itself doesn't reference or favor xAI/Grok; the Daily Driver and Signal sections if anything lean toward Anthropic/OpenAI models (Fable, GPT-5.6 Sol, Opus 5), so the ad reads as incidental inventory rather than editorial influence. Also present, per Every's recurring pattern: house self-promo for the Every All Access bundle (Sparkle, Cora, Spiral, Monologue) and a "Want to sponsor Every?" CTA — standard, not a new bias signal.
Related
- [[2026-08-25-every-benchmarks-dont-know-your-job]]
- [[2026-08-20-every-defense-of-ai-writing]]
- [[2026-08-06-every-codex-of-ones-own]]
- [[feedback_delegation_model_effort_pairing]]
- [[feedback_implementation_notes_sub_agent_pattern]]