Why this is in the vault
The DISCUSS segment carries a named-source, on-record claim from OpenAI's Core Agent lead that harnesses should trend toward minimal scaffolding as models improve — direct outside validation for RDCO's skills-over-harness-complexity posture.
The core argument
DISCUSS: "Will harness engineering go the way of prompt engineering?" Joe Gershenson, who leads OpenAI's Core Agent team for ChatGPT Work and Codex, told Every: "The harness is very important now, but in the long run, you want to keep it as simple as possible. Models are going to get smarter, and the harness is going to get better at getting out of their way." Every's framing: agent harnesses (the scaffolding — tools, context, guardrails — around a model) went mainstream in late 2025 when frontier models got capable enough to handle tasks autonomously, forcing labs to build out prescriptive scaffolding fast. As models improve, that scaffolding is expected to get simpler, not more complex — "the high-level trend in harness engineering will be finding ways to give the model more degrees of freedom." It's a single sourced quote, not a researched thesis piece, but it's from someone with direct visibility into how OpenAI builds Codex's own harness.
Curation section
- Lead story — "The Agents Behind the Curtain": Naveen Naidu, solo founder of Monologue (Every's dictation app), runs his one-person shop as a managed team of Codex agents — each a separate Codex project with its own
AGENTS.md, skills, folders, memory, and codebase access (customer support agent, web engineer agent, growth strategist agent). GPT-5.6 now lets one project hand off context and kick off a task in another project directly ("Sent by Codex from another chat"), which Naveen uses for cross-agent handoffs, e.g., support agent → web engineer agent to add a customer testimonial. - Steal This Workflow — dispatch desk: Naveen's triage pattern — keep one thread per project as the intake point, decide what the agent handles itself vs. routes out, and use a fixed handoff template ("Review this issue, create a new worktree in [project] to [complete the task], and [produce the deliverable]") to pass specialized work to the right agent.
- AI & I podcast (re-air): Portola/Tolan — an AI "alien companion" app pulling in ~$4M/year. Cofounder Quinten Farmer (sold a prior fintech company for $300M) and head of story Eliot Peper (sci-fi novelist) argue LLMs are a new storytelling medium, not just a tool. Tolan is deliberately not scripted — trained to improvise like an actor rather than follow a narrative outline — and onboarding uses a personality quiz to make the companion feel "familiar," not identical. Their bet: consumer AI repeats the automobile's arc from Model-T utility (ChatGPT) to identity-expressive products, ending in "character-driven computing" where a trusted character, not a search bar, is the first AI touchpoint.
- Tech Stack — "Thesis 2027 Edition": How Every's design team built the brand identity, website, and launch assets for its own conference (Thesis) in ~3 weeks vs. an estimated 4+ months pre-AI. Workflow: Pinterest/Cosmos for mood-boarding → Figma for a locked brand system before any web work started → Midjourney + ChatGPT Images 2.0 for visual assets → Unicorn Studio for motion overlays → Claude (Opus 4.8) to turn approved copy/UX flows into rough wireframes → Figma again for final hand-built layouts using Claude's wireframes as reference. Notable discipline: lead designer refused to start the website until the brand system was locked ("It has to make sense visually first").
Mapping against Ray Data Co
Gershenson's quote is direct outside confirmation of the bet already encoded in the "Skills over commands" memory (~/.claude/skills/ format, never ~/.claude/commands/) and the harness-thesis-dissent cluster this vault already tracks (see Related): if the labs themselves are actively simplifying their own harnesses and pushing degrees of freedom back to the model, RDCO's differentiation can't live in harness cleverness — it has to live in skill depth and domain judgment, which is exactly the target of the /improve skill's self-improvement loop. This is evidence for that bet, not neutral color.
Secondary, weaker mapping: Naveen's per-specialist Codex-project setup (dedicated AGENTS.md, dispatch-desk triage, cross-project handoff) is an independent, real-world confirmation of the same shape as RDCO's own brigade pattern (station-spec-author → station-test-author → station-code-author → station-critic). The difference is structural — Naveen's setup is ad hoc and persona-based, RDCO's is a fixed 4-station pipeline — so this is a "someone else converged on something adjacent" data point, not a validation of the specific pipeline design.
The Thesis build workflow (Pinterest/Cosmos → Figma brand lock → Midjourney/ChatGPT Images → Claude wireframes → Figma final) is a concrete, recent example of an AI-assisted brand/website production pipeline; worth a glance if ray-data-co-design or build-landing-page ever need a reference for how a comparable shop sequences brand-lock-before-build.
⚠️ Sponsorship
Not third-party sponsored. The issue closes with Every's standard house cross-promotion: an "Every All Access" upsell (bundles membership + a "$9,000+ Builder Pack" of tool credits) and a footer "Bundle of AI software" pushing Every's own products (Sparkle, Cora, Spiral, Monologue — the last of which is also the subject company in the lead story). This is boilerplate footer self-promo, not woven into the editorial content itself, and Monologue's presence as the lead-story subject is disclosed implicitly by it being Every's own product — worth noting as a house-promotion pattern but it doesn't materially bias the Gershenson quote or the Thesis workflow writeup, which are the two load-bearing pieces here.
Related
- [[2026-04-11-garry-tan-thin-harness-fat-skills]]
- [[2026-04-13-every-folder-is-the-agent]]