06-reference

3blue1brown last imo problem ai could not solve

2026-09-18·reference·source: 3Blue1Brown (YouTube)·by Grant Sanderson
AImathematicsIMOreasoning-modelsagent-capability

"The last IMO problem AI could not solve" — 3Blue1Brown

Why this is in the vault

A concrete, dated data point on the frontier of AI reasoning capability (2024 AlphaProof: 4/6 IMO problems with Lean assist → 2025: 5/6 in natural language → 2026: 6/6 via public reasoning models) plus a first-person account from a Google DeepMind research director on why models still fail on a specific class of problem — directly relevant to RDCO's agent-capability-driven thesis.

Episode summary

Grant Sanderson walks through the hardest problem (P6) on the 2025 International Math Olympiad — a tiling-and-optimization puzzle that no AI model solved that year, though models cleared it by 2026. He builds the full construction and proof from first principles (windmill tiling, Erdős–Szekeres theorem), then reflects on why this class of problem was hard for AI and argues for elevating "motivated explanation" (the narrative that makes a proof feel inevitable) to the same status as proof itself in mathematical practice.

Key arguments / segments

Notable claims

Mapping against Ray Data Co

This is a direct empirical touchpoint for RDCO's core thesis that bets are downstream of agent capability (see project_l5_north_star_strategic_direction). The specific failure mode identified — models rushing to solve before "understanding" a problem, and lacking patience/taste for elegant strategies — is a generalizable diagnostic for where agentic systems still underperform on ill-specified, multi-step reasoning tasks, not just contest math. It's a useful benchmark data point to track alongside phData's Anthropic/Snowflake cert work (project_phdata_cert_escalator_path) and any future claims about "AI can/can't do X yet" — the IMO went from AGI-adjacent litmus test (2023) to fully solved (2026) in three years, which should recalibrate any capability-ceiling assumptions baked into RDCO's roadmap. Also relevant to Ray's own operating pattern: the "patience before solving" failure mode is worth watching for in Ray's own agent dispatches (verify-dispatch, PRE-DECOMP discipline already guard against premature execution).

Related