"The Model Is the Easy Part" — Every Staff (Alex Duffy; Ashwin Sharma)
Why this is in the vault
Two original essays — Duffy on measurement-as-moat for AI adoption and Sharma on why standalone AI drug discovery won't create durable value — plus a Knowledge Base digest of articles largely already filed individually; keeping this for the Duffy thesis which has direct bearing on the phData CAF layer.
⚠️ Sponsorship
Main essay "MEASURE WHAT MATTERS—AND GET PAID FOR IT" by Alex Duffy is authored in the first-person plural of Good Start Labs ("Over the past year at Good Start Labs, we've built benchmarks…"). No explicit "Sponsored" disclosure appears in the email, but the essay functions as native advertising — it names a client case study (Arkadium / Game Lab) and describes Good Start Labs' service offering in detail. Treat the Duffy section as commercially motivated content. "FROM EVERY STUDIO" block is standard Every self-promo (Every All Access, Builder Pack, Spiral, Monologue, Cora).
The core arguments
Essay 1: "The Model Is the Easy Part" — Alex Duffy (Good Start Labs)
The argument is that picking a frontier model is the solved problem; defining what "good" means for your specific use case and then measuring toward it is where enterprise AI value gets created and retained. Two payoffs emerge from rigorous goal-definition: (1) better internal ROI because you can tell valuable AI use from noise, and (2) revenue from selling well-labeled domain data to frontier labs hungry to improve in non-mainstream tasks.
The case study: Arkadium has millions of casual gamers. Good Start Labs helped them discover that frontier LLMs lose at Gin Rummy ~90% of the time against casual players — because those models have seen lots of math and almost no card-game play. The fix wasn't a bigger LLM; it was a 4.6M-parameter expert model in an 18MB file, running on a CPU at roughly $60/year at 1M daily requests vs. multi-million-dollar LLM costs. Arkadium then monetized its well-structured gameplay data to frontier labs for millions in recurring revenue. Duffy cites Reddit, Shutterstock, and News Corp as others who have captured "hundreds of millions" in lab data deals.
Key caveat: companies whose data IS their product must protect it. Duffy names Figma/Anthropic and Cursor/Claude as cautionary examples where the lab relationship turned adversarial. Anthropic asked pharma for data; nearly everyone said no. The defensibility of the data market long-term is unknown — but the essay's closing bet is that "defining and measuring 'good' is emerging as the next stage of AI adoption."
Essay 2: "The Narrow Promise" (Alignment section) — Ashwin Sharma
AI drug discovery companies are solving only discovery-to-lead — the dark-blue sliver at the far left of the pipeline. Everything after it (assay/cell-line validation, manufacturing, animal tox, three phases of human trials) is ~99% of the remaining bar. If models commoditize, standalone AI-discovery companies may be worth no more than a ChatGPT enterprise subscription with a neat pitch deck. Full-stack biotechs that own lab work AND AI will capture the value. The abstract lesson: "The future may look less like software eating pharma than pharma eating software." Computational abundance makes downstream skilled judgment more scarce and valuable, not less.
Issue contents
Knowledge Base (curated articles)
Most of these have already been individually filed:
- "How I Polish Software That Agents Built" — Kieran Klaassen / Source Code → see [[2026-07-13-every-polish-agent-built-software]]
- "The Case Against Skills" — Laura Entis / Mike Taylor / Context Window → see [[2026-07-16-every-case-against-skills]]; argues frontier models have absorbed most trending skills, piling instructions degrades outputs
- "The Urge to Merge (ChatGPT and Codex)" — Katie Parrott / Context Window → see [[2026-07-14-every-chatgpt-codex-merge-revolt]]; OpenAI folded Codex into ChatGPT desktop, power-user revolt
- "The Ops Team That Routes Work Across Models" — Laura Entis / Context Window → see [[2026-07-15-every-biz-ops-ai-workflows]]; Every's biz-ops team using Fable + Codex + Fin in fluid cross-model workflows
- "The Founder of a $1.5 Billion AI Company on What Comes After the First Wave of AI Apps" — Dan Shipper / AI & I podcast → Granola CEO Chris Pedregal; Granola watched Notion/OpenAI/Zoom copy its meeting-notes feature; betting on owning work around meetings, not just notes
- "I'm an Editor—And I Built Our Newest Feature" — Jack Cheng / On Every → see [[2026-07-17-every-gift-links-ai-native-build]]; senior editor (non-engineer) built gift-link feature with Codex
From Every Studio (house promos)
Every All Access annual membership launched with a Builder Pack (~$7K in tool credits). Sandbar's Stream is first external product built on Monologue's voice API. Spiral expanded writing rules to support structural restructuring and multi-language drafts.
Mapping against Ray Data Co
The Duffy measurement thesis maps directly onto Ray's phData CAF PM role. The CAF "Fabric" governed knowledge graph is precisely the instrumentation and measurement layer that Duffy argues is the next enterprise AI battleground — it's the layer that defines what "good" looks like for an organization's AI outputs and tracks progress against it. The essay gives a practitioner-level vocabulary (benchmark construction, domain-specific expert models, data monetization to frontier labs) that can be surfaced in phData client discovery conversations to position the CAF layer as strategic, not just architectural.
The Sharma drug-discovery essay extends to a broader RDCO belief already held: AI commoditizes the early-stage, high-throughput part of any pipeline; durable advantage lives in owning the judgment-intensive downstream. This is the same logic Ray applies to RDCO's advisor positioning — the model is not the moat, the instrumented workflow and domain judgment are.
The "Case Against Skills" thread (see [[2026-07-16-every-case-against-skills]]) raised the question of whether RDCO's growing ~/.claude/skills/ library is becoming overhead as frontier models absorb those capabilities. This digest's Knowledge Base section continues that thread (Mike Taylor argues exactly this; Naveen Naidu dropped all but one skill). Worth a periodic review pass on the skills library to prune redundant instruction stacks.
Related
- [[2026-07-16-every-case-against-skills]]
- [[2026-05-22-benn-stancil-wac-wins-above-claude]]
- [[2026-05-11-cfo-secrets-ai-for-cfos-series-synthesis]]