"OpenAI Does Math, Reward-Hacking, Meta Launches Personal Agent" — Ben Thompson
Why this is in the vault
Two threads in one Update both land on RDCO's own build: Thompson's reward-hacking framing is the exact failure mode RDCO's fresh-eyes critic architecture exists to catch, and Meta's Muse launch is a direct competitive/comparative data point for the Channels agent (Mac Mini always-on Discord/iMessage agent) and the L5 agent-capability bet.
Mapping against Ray Data Co
The load-bearing connection is Thompson's read of the OpenAI Navier-Stokes controversy: mathematician Terence Tao's critique (quoted at length) is that AI tools can "flatten" a problem space and get credited with solving it while destroying the actual research value — and Thompson extends this to argue OpenAI itself is reward-hacking, nominally achieving "we solved a hard math problem" while gaining nothing (no new technique, no new understanding). This is precisely the failure mode named in [[feedback_workflow_agent_output_integrity]] (pointer returns, false "verified" stamps, proposal-as-fact) — an agent optimizing for the appearance of a completed goal rather than the goal itself. Thompson's line that "humans provide the why of doing something; absent that an AI can do nothing more than obsess over proxies" is a one-sentence restatement of why RDCO gates strategic outputs and vault writes through an independent fresh-eyes critic ([[feedback_verification_independent_worker_pattern]]) rather than trusting a producer agent's self-report.
Separately, Meta's Muse — a WhatsApp-native personal agent backed by a persistent per-user VM (8GB RAM/disk), free tier plus $20/$100 subscription tiers, that "remembers what matters" and acts on standing goals — is the first mainstream-consumer instantiation of the always-on personal agent RDCO has been running internally via the Channels agent ([[project_channels_agent_setup]], Mac Mini + LaunchAgent + tmux). Thompson's verdict that Muse's launch matters more to "most of humanity" than OpenAI's math result than any prestige benchmark is a direct data point for RDCO's L5 north star thesis that bets are downstream of agent capability, not model capability ([[project_l5_north_star_strategic_direction]]) — a well-known consumer platform is now shipping the exact wrapper (persistent memory, proactive suggestion, approval-gated purchases/emails) RDCO has been building bespoke, which both validates the shape and raises the competitive bar on distribution.
Issue contents
OpenAI Does Math / Reward-Hacking
OpenAI reportedly solved one of the six Millennium Prize ("million-dollar") math problems — proving flaws in the Navier-Stokes fluid-motion equations — via a single LLM prompt. Mathematician Tristan Buckmaster (with Anthropic's Levent Alpöge) alleges OpenAI got wind of their unpublished progress and used it to beat them to the result, possibly via chat-log signal; OpenAI denies this. Thompson is agnostic on the espionage allegation (he notes ChatGPT's opt-out training settings have real loopholes — thumbs-up feedback still trains on the full conversation) but foregrounds Terence Tao's structural critique: AI tools flatten a field's difficulty landscape and reward-hack the appearance of progress without producing new technique or understanding. Thompson ties this back to his own prior-day Article "Write Things Down": absent a human-supplied "why," an AI (or an organization) can only optimize proxies.
Meta Launches Personal Agent (Muse)
Meta launched Muse, a free (with $20/$100 paid tiers) personal AI agent accessible via a dedicated app and WhatsApp, running on a dedicated per-user "Muse Secure VM" and powered by Meta's new "Muse Spark" model. It handles tasks (email, bookings, purchases with approval gates), remembers context unprompted, and proactively suggests actions. Thompson, who has early access, reports it works well and closes real gaps (reminders, idea tracking) with a much lower setup cost than his own bespoke multi-Mac-Mini agent fleet. His verdict: Muse matters more than OpenAI's math result because it's aimed at "human scale, with computers, instead of computer scale, in place of humans" — mainstream consumer adoption of agents is the real open question, not frontier benchmark performance.
Related
- [[2026-09-08-stratechery-write-things-down-harness-memory]]
- [[2026-07-09-stratechery-update-muse-grok-karp]]
- [[feedback_workflow_agent_output_integrity]]
- [[feedback_verification_independent_worker_pattern]]
- [[project_channels_agent_setup]]
- [[project_l5_north_star_strategic_direction]]