Nobody Else Has Shipped This: The Public-Channel Agent Is an N-of-1 With an N-of-0 Control
The question
"Beyond Shopify's River, which companies have shipped agents that refuse DMs and force public channels — and what's the measured productivity / learning delta vs DM-mode agents?"
Extends the 2026-05-09 Tobi Lütke River piece (36% → 77% merge rate via public-channel osmosis). Load-bearing for Ray harness positioning + MAC executable-course delivery design.
What we already know (from the vault)
- [[2026-05-09-tobi-lutke-river-public-channel-agent]] is the vault's only case study, and it is a primary source: Tobi Lütke's own X long-form article. Reports River refusing DMs by design, 5,938 employees across 4,450 channels in 30 days, 1,870 PRs in one week, ~1 in 8 merged monorepo PRs River-authored, and the headline merge rate 36% → 77% over two months with no model retrain and no model switch.
- The same note already flagged the mechanism as org-layer harness engineering: "model improves slowly, harness/skill/instruction layer improves daily because the entire company is watching where it fails."
- The note also pre-registered the exact skepticism this brief now tests: "Tobi reports 5,938 employees — does it work at 50? 5? 1+AI?" That question remains unanswered by any source found here.
- [[2026-05-10-agent-harness-landscape]] positions River as one of ~6 independent operators converging on harness-engineering, but cites no second public-channel implementation.
- [[2026-06-04-technically-sentry-non-engineers-ship-cms-to-git]] is the closest vault analog to an org-learning delta (~2,500 pages, 2.5 developers, two months, Claude Code writing most code) — but Sentry's mechanism was artifact visibility via git, not channel-visibility, and it reports no public-vs-private comparison.
What the web says
- No second company was found. Three independent sources were checked and none names another org that has shipped an agent refusing DMs. The ZenML LLMOps case study explicitly contrasts River against GitHub Copilot (private IDE suggestions), Cursor (private chat), Claude, ChatGPT, and Devin — and states that no competing implementation with mandatory public channels is identified.
- Shopify's own engineering blog does not corroborate the 36→77% figure. Under the River reports a later, different 30-day window — 59,918 River sessions, 5,170 channels, 7,000+ people touched, 3,536 River-coauthored PRs merged, median session 19 minutes, median 50 tool calls per session — all self-reported from the
river_sessionsdomain table. It reports no merge rate at all. - Under the River contains no public-vs-DM comparison. The design is justified philosophically, not experimentally: "If every interaction with an agent happens in a private window, the only person who learns anything is the person at the keyboard." No alternative implementation data is presented.
- The ZenML figures are derivative, not independent. Its "5,938 employees / 4,450 channels / one in eight PRs / 36% → 77%" match Tobi's X article exactly. This is a restatement of the vault's existing primary source, not a second measurement.
- Commentary sources add analogy, not evidence. Stack Archive's piece offers only aspirational claims ("the public model may produce compounding returns that private deployment cannot") with no quantified metrics or controlled comparison. Its one adjacent precedent is Simon Willison's parallel to Midjourney's early Discord architecture — public-by-default generation as a historical analogy, not a shipped agent that refuses DMs.
- Searches for a controlled public-vs-DM study returned nothing. General 2026 agent-productivity literature exists (Microsoft Work Trend Index; arXiv 2503.18238 on human-AI teamwork; the Microsoft Claude Code / Copilot CLI rollout study at arXiv 2607.01418), but none isolates channel visibility as the variable. "Public vs private AI" in that literature means cloud-vs-on-prem infrastructure, not social visibility.
Convergences and contradictions
- Convergence (strong): Vault and web agree on the design constraint and on River's scale. Every source agrees River is the only shipped instance of the pattern. The negative finding is robust across three independent checks.
- Contradiction (important): The vault treats "merge rate doubling" as the load-bearing data point. The web reveals it rests on a single self-reported figure from one CEO's marketing-adjacent essay, never repeated in Shopify's own engineering deep-dive on the same system. Under the River had every opportunity to restate 36→77% and did not.
- Contradiction (framing): The vault note called River "a direct empirical validation of the architecture pattern." That overstates it. River is an uncontrolled single-arm observation — no DM-mode control group, no counterfactual, self-reported telemetry, confounded by two months of concurrent skill-writing, prompt tuning, and model-vendor improvements that Tobi's "no retrain" claim does not rule out.
Synthesis for RDCO
The honest headline: we found no measured delta, because no one has measured one. Not a weak delta, not a contested delta — no study, anywhere, isolates public-channel vs DM-mode as a variable. And the one number we do have (36% → 77%) is single-sourced, self-reported, and conspicuously absent from Shopify's own engineering write-up of the same system. This should downgrade how load-bearing that figure is allowed to be in any RDCO artifact. It is a compelling anecdote from a credible operator, not evidence. If a Sanity Check piece or a MAC module leans on "36→77%" as proof, we are doing the thing we criticize other people for.
That said, the negative finding is more useful to RDCO than a confirmation would have been, and it cuts two ways. First, for Ray harness positioning: the pattern is uncontested territory. One company has shipped it, no vendor has productized it, and the entire competitive set (Copilot, Cursor, Devin, Claude in-IDE) is architecturally committed to the private window. That is a genuine seam — the same shape as the CAF "claim the UNOWNED wedge" move. But the wedge is not "public channels make agents better" (unproven). The wedge is "nobody can tell you whether it does, and everyone has already bet against it." RDCO can be the party that instruments the question rather than the party that repeats Tobi's number louder.
Second, for MAC executable-course delivery: this is where the finding actually bites. River's causal story is osmosis — juniors scroll back through a senior's #their_river channel and learn to scope requests by watching. That mechanism requires a population. The vault note already caught this ("does it work at 50? 5? 1+AI?") and nobody has answered it. A course product is a plausible venue to manufacture the population River gets for free from 7,000 employees: an executable course where learner-agent transcripts are public-by-default to the cohort is structurally the same Lehrwerkstatt, at cohort scale. That is a real design thesis. It is also, notably, an untested one — and MAC is the one surface where we could actually run the experiment, because we'd control both arms.
Which suggests the sharpest move available: stop citing the delta and go measure it. A cohort with public-by-default transcripts vs a cohort with private ones, same curriculum, measuring time-to-competence, is a study that does not currently exist in the literature. That is a publishable, differentiating asset — and it converts our current position (one derivative case study everyone else also read) into primary evidence nobody else has. The asymmetry is favorable: the study is cheap if MAC ships anyway, and the negative result ("public channels don't help at cohort scale") would be nearly as valuable as the positive one, since it would spare us from building an architecture on a single unreplicated anecdote.
Concretely, this argues for keeping the [[2026-05-09-tobi-lutke-river-public-channel-agent]] Sanity Check candidate alive but re-framing the angle. The original working title was "The agent that refuses to work in private." The better, more honest, and more original piece is now available: the most-cited number in the public-channel-agent thesis has never been replicated — including by the company that produced it. That is a re-frame, not a restatement, and it clears the no-derivative-pieces bar.
Why this is in the vault
This brief downgrades the evidentiary weight of the 36→77% figure that [[2026-05-09-tobi-lutke-river-public-channel-agent]] treats as "direct empirical validation" — a correction that directly gates how Ray harness positioning and MAC course-design arguments are allowed to cite it. It also establishes that the public-vs-DM delta is unmeasured territory, which converts MAC from a consumer of Tobi's anecdote into a candidate venue for the first controlled test.
Open follow-ups
- Does the public-channel constraint pay back below ~1,000 people? Nobody has data. What is the minimum population where osmosis beats the coordination cost? (Carried forward unanswered from the 2026-05-09 note.)
- Did Shopify's engineering team omit the merge-rate figure from Under the River deliberately? Worth checking whether any Shopify engineer has restated 36→77% independently of Tobi.
- Does arXiv 2607.01418 (Microsoft's Claude Code / Copilot CLI rollout study) contain any visibility-related variable? Not fetched here — flagged as a lead only, no findings cited.
- What is the confound magnitude? Over the two months of River's 36→77% climb, how much is attributable to frontier-model improvements shipped by the vendor rather than Shopify's skill-writing?
- Are there non-agent precedents with real measured deltas — e.g. public-by-default code review, open-by-default incident channels, GitLab/Automattic radical-transparency orgs? The mechanism might be validated in a pre-AI literature we haven't searched.
- What is the counter-evidence? No source found names a single critique of public-only design (chilling effects, performance anxiety, sensitive-topic leakage). The absence of dissent in the corpus is itself suspicious and worth a dedicated adversarial search.
Related
- [[2026-05-09-tobi-lutke-river-public-channel-agent]] — the vault's primary and only River case study; this brief materially qualifies its "empirical validation" framing
- [[2026-05-10-agent-harness-landscape]] — positions River within the harness-engineering thesis cluster; no second public-channel instance found there either
- [[2026-04-19-acquired-tobi-lutke-shopify]] — prior Tobi entry; useful for calibrating him as a source (constitutions, trust batteries, private evals)
- [[2026-06-04-technically-sentry-non-engineers-ship-cms-to-git]] — nearest vault analog for an org-learning delta, via artifact visibility rather than channel visibility
- [[2026-05-08-jaya-gupta-shape-as-moat]] — org-shape-as-moat thesis; River is the worked example
- [[2026-05-10-addy-osmani-agent-harness-engineering]] — harness-engineering anchor doc for the surrounding cluster
Sources
Vault
~/rdco-vault/06-reference/2026-05-09-tobi-lutke-river-public-channel-agent.md→ [[2026-05-09-tobi-lutke-river-public-channel-agent]]~/rdco-vault/06-reference/research/2026-05-10-agent-harness-landscape.md→ [[2026-05-10-agent-harness-landscape]]~/rdco-vault/06-reference/2026-06-04-technically-sentry-non-engineers-ship-cms-to-git.md→ [[2026-06-04-technically-sentry-non-engineers-ship-cms-to-git]]
Web
- Shopify Engineering, "Under the River" — https://shopify.engineering/under-the-river (59,918 sessions / 5,170 channels / 7,000+ people / 3,536 merged PRs, 30-day window; self-reported via
river_sessions; no merge rate, no DM comparison) - ZenML LLMOps Database, "Building a Public AI Agent Workspace for Organizational Learning" — https://www.zenml.io/llmops-database/building-a-public-ai-agent-workspace-for-organizational-learning (restates Tobi's figures; explicitly finds no competing public-only implementation)
- Stack Archive, "Shopify's River Won't Talk to You in Private — That's the Point" — https://stack-archive.com/blog/shopify-river-agent-public-slack-2026/ (opinion; no measured data; Midjourney/Discord analogy)
- Tobi Lütke, "Learning on the Shop floor" — https://x.com/i/article/2052738533111013380 (ultimate and only source of the 36% → 77% figure)
- Simon Willison, "Learning on the Shop floor" — https://simonwillison.net/2026/May/11/learning-on-the-shop-floor/ (surfaced in search; source of the Midjourney-Discord parallel)
Searched and found nothing: controlled studies isolating public-channel vs DM-mode agent deployment. General 2026 agent-productivity literature (Microsoft Work Trend Index, arXiv 2503.18238, arXiv 2607.01418) does not treat channel visibility as a variable; "public vs private AI" in that corpus refers to infrastructure, not social visibility.