"The last IMO problem AI could not solve" — 3Blue1Brown
Why this is in the vault
A concrete, dated data point on the frontier of AI reasoning capability (2024 AlphaProof: 4/6 IMO problems with Lean assist → 2025: 5/6 in natural language → 2026: 6/6 via public reasoning models) plus a first-person account from a Google DeepMind research director on why models still fail on a specific class of problem — directly relevant to RDCO's agent-capability-driven thesis.
Episode summary
Grant Sanderson walks through the hardest problem (P6) on the 2025 International Math Olympiad — a tiling-and-optimization puzzle that no AI model solved that year, though models cleared it by 2026. He builds the full construction and proof from first principles (windmill tiling, Erdős–Szekeres theorem), then reflects on why this class of problem was hard for AI and argues for elevating "motivated explanation" (the narrative that makes a proof feel inevitable) to the same status as proof itself in mathematical practice.
Key arguments / segments
- [00:01:01] Capability timeline: 2023 Dwarkesh Patel podcast question ("IMO gold = AGI?") → 2024 AlphaProof (4/6, Lean-assisted) → 2025 (5/6, natural language) → 2026 (6/6, public models via simple prompting)
- [00:03:00] Google DeepMind's Tang Luong: models "didn't take the time to understand the problem" before attempting a solution — lack of patience, not lack of raw capability
- [00:04:01] The problem itself: 2025×2025 grid, minimum tiles such that each row/column has exactly one uncovered square
- [00:07:01] Analogous warm-up puzzle (cube-slicing, answer=6) used to motivate a 1:1 correspondence proof technique
- [00:14:00] Optimal construction: windmill pattern of k×k square tiles, k²+2k-3 total tiles for a k²-side grid; k=45 for 2025 = 45²
- [00:25:00]-[00:34:01] Core proof technique: two paths (longest increasing/decreasing subsequences of the X-position permutation) divide the grid into four regions, enabling a tight lower-bound argument
- [00:39:00] Erdős–Szekeres theorem proof via lattice-point argument, combined with AM-GM inequality to close the bound
- [00:46:01]-[00:49:00] Grant's reflection: why visual/intuition-heavy problems requiring "understanding before solving" resist reinforcement-learning training; his broader "motivated explanation" thesis (also posted as a guest piece on Terence Tao's blog)
Notable claims
- Less than 1% of IMO 2025 contestants got full marks on P6, and it was the only problem of six that all AI systems failed that year
- By 2026, "simply prompting reasoning models that are available to the public would be enough to get correct answers on all six problems" — a one-year capability jump on the hardest available math benchmark
- Grant's framing: "it warmed my heart when I solved it" (a competitor's own words, quoted from a Patreon comment) is offered as a distinct, human, non-proof source of mathematical value
Mapping against Ray Data Co
This is a direct empirical touchpoint for RDCO's core thesis that bets are downstream of agent capability (see project_l5_north_star_strategic_direction). The specific failure mode identified — models rushing to solve before "understanding" a problem, and lacking patience/taste for elegant strategies — is a generalizable diagnostic for where agentic systems still underperform on ill-specified, multi-step reasoning tasks, not just contest math. It's a useful benchmark data point to track alongside phData's Anthropic/Snowflake cert work (project_phdata_cert_escalator_path) and any future claims about "AI can/can't do X yet" — the IMO went from AGI-adjacent litmus test (2023) to fully solved (2026) in three years, which should recalibrate any capability-ceiling assumptions baked into RDCO's roadmap. Also relevant to Ray's own operating pattern: the "patience before solving" failure mode is worth watching for in Ray's own agent dispatches (verify-dispatch, PRE-DECOMP discipline already guard against premature execution).
Related
- [[2026-06-30-dwarkesh-grant-sanderson-3blue1brown-ai-future-math]]
- [[2026-04-20-3blue1brown-imo-geometry-alphageometry-aleph0]]
- [[2026-03-20-dwarkesh-patel-terence-tao-how-the-worlds-top-mathematician-uses-ai]]
- [[project_l5_north_star_strategic_direction]]