06-reference

every engineering team cost of codex

2026-08-19·reference·source: Every (Context Window)·by Laura Entis
harness-engineeringagent-architecturecodexskillsai-design-workfloweverythin-harness-fat-skills

Why this is in the vault

The DISCUSS segment carries a named-source, on-record claim from OpenAI's Core Agent lead that harnesses should trend toward minimal scaffolding as models improve — direct outside validation for RDCO's skills-over-harness-complexity posture.

The core argument

DISCUSS: "Will harness engineering go the way of prompt engineering?" Joe Gershenson, who leads OpenAI's Core Agent team for ChatGPT Work and Codex, told Every: "The harness is very important now, but in the long run, you want to keep it as simple as possible. Models are going to get smarter, and the harness is going to get better at getting out of their way." Every's framing: agent harnesses (the scaffolding — tools, context, guardrails — around a model) went mainstream in late 2025 when frontier models got capable enough to handle tasks autonomously, forcing labs to build out prescriptive scaffolding fast. As models improve, that scaffolding is expected to get simpler, not more complex — "the high-level trend in harness engineering will be finding ways to give the model more degrees of freedom." It's a single sourced quote, not a researched thesis piece, but it's from someone with direct visibility into how OpenAI builds Codex's own harness.

Curation section

Mapping against Ray Data Co

Gershenson's quote is direct outside confirmation of the bet already encoded in the "Skills over commands" memory (~/.claude/skills/ format, never ~/.claude/commands/) and the harness-thesis-dissent cluster this vault already tracks (see Related): if the labs themselves are actively simplifying their own harnesses and pushing degrees of freedom back to the model, RDCO's differentiation can't live in harness cleverness — it has to live in skill depth and domain judgment, which is exactly the target of the /improve skill's self-improvement loop. This is evidence for that bet, not neutral color.

Secondary, weaker mapping: Naveen's per-specialist Codex-project setup (dedicated AGENTS.md, dispatch-desk triage, cross-project handoff) is an independent, real-world confirmation of the same shape as RDCO's own brigade pattern (station-spec-author → station-test-author → station-code-author → station-critic). The difference is structural — Naveen's setup is ad hoc and persona-based, RDCO's is a fixed 4-station pipeline — so this is a "someone else converged on something adjacent" data point, not a validation of the specific pipeline design.

The Thesis build workflow (Pinterest/Cosmos → Figma brand lock → Midjourney/ChatGPT Images → Claude wireframes → Figma final) is a concrete, recent example of an AI-assisted brand/website production pipeline; worth a glance if ray-data-co-design or build-landing-page ever need a reference for how a comparable shop sequences brand-lock-before-build.

⚠️ Sponsorship

Not third-party sponsored. The issue closes with Every's standard house cross-promotion: an "Every All Access" upsell (bundles membership + a "$9,000+ Builder Pack" of tool credits) and a footer "Bundle of AI software" pushing Every's own products (Sparkle, Cora, Spiral, Monologue — the last of which is also the subject company in the lead story). This is boilerplate footer self-promo, not woven into the editorial content itself, and Monologue's presence as the lead-story subject is disclosed implicitly by it being Every's own product — worth noting as a house-promotion pattern but it doesn't materially bias the Gershenson quote or the Thesis workflow writeup, which are the two load-bearing pieces here.

Related