Twitter/X

vLLM runs on roughly 500,000 GPUs at any given moment, per Simon Mo; Simon Mo is…

Brief

vLLM is an open-source inference project that, according to lead maintainer Simon Mo (Inferact CEO), operates on about 500,000 GPUs at any given moment. In an a16z interview (Aug 6, 2026) Mo outlines the operational realities of running open models in production—day‑zero releases, Kimi K3 tradeoffs, licensing shifts, funding analogies, and pre‑ChatGPT origins.

Why it matters

vLLM runs on roughly 500,000 GPUs at any given moment, per Simon Mo; Simon Mo is co-founder and CEO of Inferact and the lead maintainer of vLLM.

Key details

  • The a16z interview (published 2026-08-06) with Matt Bornstein and Elena Burger catalogs production topics including day-zero model releases, what the Kimi K3 hardware actually buys you, and the fact that vLLM, OpenRouter, and Ollama all started before ChatGPT.
  • Simon details shifting open-model licenses, uses a pharmaceutical analogy for funding model training, explores a hypothetical 99% GPU price drop, and notes that the inventor of RoPE removed RoPE.
Source evidence

vLLM runs on half a million GPUs at any given moment. Most people have never heard of it.

Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more.

00:00 Intro
01:46 When open source became critical infrastructure
08:55 Day zero model releases, and the drama behind them
14:59 What Kimi K3 actually buys you
18:56 Why open model licenses are changing
22:24 The pharmaceutical analogy for funding model training
26:16 If GPUs got 99% cheaper
29:08 Why vLLM, OpenRouter, and Ollama all started before ChatGPT
35:42 Building a company on an open source project
39:48 The inventor of RoPE removing RoPE

@simonmo @BornsteinMatt @VirtualElena

Video