06-reference/transcripts

dwarkesh why smarter ai models could drive up compute prices 10x transcript

2026-08-03

Raw transcript — "Why smarter AI models could drive up compute prices 10x" (Dwarkesh Patel)

Today I want to talk about what the compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropic's revenue has 10xed year-over-year and it's likely to do so again this year. So they ended last year with 9 billion in revenue. I think they'll probably end this year with somewhere between 100 billion to $150 billion in revenue. Now for this trend to continue, Anthropic would need to make $1 trillion in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities. Does AI get that useful by the end of next year? But suppose the trend does continue. Well, I want to think through what happens in that world.

[00:01:00] Now, the other big trend in AI is that lab compute only 3xes year-over-year. For a lab to keep 10xing revenue year-over-year while compute only 3xes, one of three things needs to happen or some combination: (1) lab margins have to increase, (2) the price of compute has to increase, or (3) the percentage of compute that labs spend on inference rather than training has to increase. My understanding is basically all three are already happening. Anthropic's inference margins reportedly went from 40% in the middle of last year to upwards of 80% now. Spot prices for compute are more than 40% higher than they were in the February trough this year. And the share of compute going to inference vs training — in 2024 per Epoch, OpenAI spent just a quarter of its compute on inference, and that number is likely closer to 50% or higher now.

[00:02:00] Labs would prefer not to shift more compute to inference — the whole point of inference revenue is to convince investors to fund the next bigger, better model. If you're spending most compute on inference, you're implicitly declaring AI progress has stalled and you're just a cloud provider — a less compelling business than building AGI. Labs think within a year they'll build models that make current ones look extremely bad, so they need to keep most compute on training/experiments. That leaves two levers: lab margins increase (labs capture surplus) or price of compute increases (everyone below the lab in the stack captures surplus).

[00:03:01] Margins already went from ~80% to potentially >90% for some top models — but that would require competitive moats that seem implausible for "intelligence" to sustain (things get competed away). So the more likely escape valve is rising compute prices. Case study: Google and Anthropic renting compute from CoreWeave/hyperscalers — Google is paying $900 million/month for 110,000 GPUs (blend of GB200s/GB300s), at 2x the spot price per hour for those GPUs, and that spot price is itself >40% higher than in February.

[00:04:00] Key conclusion: as AI models get smarter, they'll better monetize the same compute. If a true human-level software engineer could run on an H100-equivalent, at today's software engineer prices that H100 should rent for over $250K/year — 15x+ the current spot price for an H100 (and that's before accounting for AI working nights/weekends). Counter-consideration: 10 million extra software engineers appearing would normally crater the marginal value of an engineer (and thus the H100) — but Dwarkesh isn't sure the lump-of-labor fallacy analogy holds here.

[00:05:00] Standard economics (e.g., high-skill immigration doesn't reduce long-run wages, per innovation/specialization effects) suggests marginal value of labor — and thus compute — could stay high. Implication: as top labs get better at monetizing compute, it becomes harder for anyone else to compete for that same scarce resource, since they're bidding against someone who can extract more value from it. Second, and most interesting implication: the Alchian-Allen effect — training the most efficient model lets you charge much higher margins, because using a weaker/less-efficient model burns more tokens on now-expensive compute for the same result.

[00:06:01] If you build a model that achieves the same result with less compute, you've effectively created more compute, and the value of compute rises. Also: many current popular AI applications will get priced out as frontier labs outbid casual usage (AI research automation) for the same tokens.

[00:07:02] Dwarkesh flags his own analysis pattern-matches to historical wrong-about-scarcity predictions (the Simon-Ehrlich bet — Paul Ehrlich's Malthusian bet on commodity prices that he lost, illustrating how market signals/ingenuity solve scarcity). He thinks that analogy is probably wrong here because compute supply is much less elastic than metal extraction and much less able to absorb demand shocks via substitution.

[00:08:00] Why the 3x/year compute growth is hard to sustain or accelerate — decomposed into three multipliers: 1.4x from Moore's Law (already miracle-level to sustain), 1.2x from new fabs (bottlenecked by ASML EUV machine production through 2030+, referencing a prior podcast guest Dylan Patel discussion), and 1.8x from AI absorbing wafer allocation previously going to smartphones/PCs — which will hit a wall by end of next year when leading-edge N3 nodes at TSMC go from 60% to 86% AI allocation, i.e., near-saturation.

[00:08:40ish] [Mercury sponsor read: Dwarkesh describes using Mercury's "Command" AI feature to auto-categorize business transactions monthly, syncing with QuickBooks. Standard fintech disclosure: "Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Column N.A., members FDIC."]

[00:10:02] Clarification: eventually compute gets cheap again — post-singularity, robots converting silica sand and copper into chips directly, at which point compute cost is just raw inputs + processing tools. This analysis is about the current "pre-singularity" regime where compute only 3xes/year, insufficient to offset AI's rising value. Anthropic's 10x revenue growth vs 3x compute growth illustrates strong economies of scale in the model business — training cost is a one-time cost amortized across all users, unlike human labor which must be retrained from scratch per instance.

[00:11:00] Closing: Dwarkesh notes he wishes economies of scale for intelligence weren't so strong, given concerns about power concentration, but believes it is the case. This is a narrated blog post also published at dwarkesh.com.