Why this is in the vault
Ben Thompson explains how Anthropic's new Claude watermarking (a green-list/red-list token-probability bias, similar in spirit to Google's SynthID) works technically and argues the EU-mandated feature is both over- and under-inclusive as a provenance signal — worth keeping as the clearest available explainer of how LLM watermarking works plus a sharp critique of AI-provenance regulation.
The core argument
Anthropic is watermarking Claude-generated text and files to comply with the EU's Code of Practice on Transparency of AI-generated Content, which requires machine-readable, detectable marking of AI outputs over 200 tokens. Thompson explains the likely mechanism: at each token-generation step the model splits candidate tokens into a keyed-hash "green list" and "red list" and slightly boosts green-list probability, producing a statistically detectable skew without changing outputs deterministically. (Google's SynthID uses a related but more complex "tournament sampling" approach designed to be less perceptible.)
Thompson's core objection is philosophical, not just technical: watermarking treats AI as an independent authorial entity that must be flagged apart from humans, when he views AI as a tool wielded by a human author — akin to demanding a pen "sign" what it writes. He also flags a practical unfairness: per Anthropic's own disclosure, a watermark can appear even when a human only used Claude to proofread or translate their own original text, mislabeling human-authored content as AI-generated. Simultaneously, the mark is trivially avoidable (heavy editing, translation, screenshots, non-EU models like SpaceXAi's or Chinese labs that never signed the EU accord) — so the regime is "too aggressive" against light AI-assisted human work and "not good enough" against anyone actually trying to hide AI origin. He ties this to his 2022 "AI Unbundling" framework: creation is the last unbundled link in the idea-propagation chain, and EU-style provenance mandates effectively reassign human creative credit to the tool.
Mapping against Ray Data Co
RDCO ships AI-assisted deliverables (Sanity Check drafts, client-facing artifacts, vault notes) that are unambiguously human-directed but pass through heavy Claude assistance at the drafting/proofing stage — exactly the case Thompson flags as mislabeled by watermarking. If RDCO or a client ever operates in the EU or under EU-adjacent disclosure norms, "was this AI-generated" becomes a binary trap: current tooling can't reliably distinguish "Claude wrote this" from "a human wrote this and had Claude proofread it," which is RDCO's actual workflow on most public output. This reinforces the existing RDCO governance posture (see the AI-governance notes below) that provenance/disclosure policy should be authored around actual human-in-the-loop process, not treated as solvable by a technical watermark — a distinction worth keeping sharp before RDCO ever makes a client-facing "AI-generated" or "AI-assisted" claim.
Related
- [[2026-04-12-lindstrom-board-ai-governance]]
- [[2026-06-15-stratechery-ben-thompson-anthropic-safety-superpower]]
- [[2026-03-02-stratechery-anthropic-and-alignment]]