Twitter/X

Two massive accelerations in the AI inference market occurred in the past year…

Brief

Author @SouthernValue95 argues the AI inference market saw two discrete accelerations—Claude Code in Jan–Mar and Fable/Codex this summer (peaking amid an AI trade implosion in July)—that pushed demand beyond supply. With 12–18 month datacenter lead times, supply shortages produced memory inflation and pricing power for clouds, while labs report rising ARR per GW of inference capacity.

Why it matters

Two massive accelerations in the AI inference market occurred in the past year: Claude Code in Jan–Mar and a second wave this summer driven by Fable/Codex, which coincided with the AI trade imploding in July.

Key details

  • AI datacenters have 12–18 month lead times, so the sudden leaps in model capability outpaced supply, driving year‑to‑date inflation across the AI supply chain—most notably memory—visible in rising Lab ARR tracked via press leaks and third‑party estimates.
  • Scarce compute has created pricing power for cloud providers and increased AI labs’ ARR per GW of inference capacity, a trend the author argues markets are beginning to appreciate but still underestimate.
Source evidence

In the past year, the AI inference market has seen two massive accelerations, the 1st was Claude Code in Jan-Mar, and we're seeing the second one right now with Fable/Codex this summer (ironically as the AI trade was imploding in July). Both accelerations are clearly visible in Lab ARR (loosely tracked via press leaks and 3P est.) (slide 1/2):

Since AI datacenters have 12-18mo lead times, supply could not keep pace with these sudden leaps in model capability and demand, driving inflation for anything in the AI supply chain YTD (most notably memory). This has created pricing power for clouds controlling scarce compute, which I've written about and the market is beginning to appreciate (but still underestimates), and for AI labs, which have seen rising ARR per GW of inference capacity (slide 2/2 next tweet):