Twitter/X

On 2026-08-06 Simon Mo (Inferact CEO, vLLM lead maintainer) said 'Open weight…

Brief

Simon Mo and Matt Bornstein argue that open-weight models are essential for cost reduction and operational control. Mo (Inferact CEO, vLLM lead) says 'Open weight absolutely matters' to avoid proprietary-API lock-in, enable customization and guardrails, and ensure predictable latency for applications like voice agents. a16z highlights vLLM running on half a million GPUs and discusses licenses, K3, and RoPE.

Why it matters

On 2026-08-06 Simon Mo (Inferact CEO, vLLM lead maintainer) said 'Open weight absolutely matters', arguing open weights prevent proprietary-API control that would block open development and research.

Key details

  • Mo identifies two drivers: control (own infrastructure, extend models, add guardrails, and guarantee latency—critical for voice-agent companies) and cost (closed models became 'too expensive'); he says control mattered over the last few years and cost became critical in the last few months.
  • a16z notes vLLM 'runs on half a million GPUs at any given moment'; the 39:48 interview covers day-zero model releases, changing open-model licenses, what Kimi K3 buys you, the RoPE inventor removing RoPE, and funding analogies for model training.
Source evidence

Simon Mo and Matt Bornstein on why open-weight models matter: cost and control.

"Open weight absolutely matters in the ecosystem. The world cannot just be controlled by proprietary APIs, where open weight, open development and research of these models are blocked or banned."

"There's almost two pieces to this. There's the cost thing, where the closed models are too expensive, and then there's the control thing, where I want to be in control of my infrastructure and in control of the model. If I need to extend it or put on my own guardrails or anything."

"Control matters a lot over the last few years, and then cost just started to matter over the last few months... It's about controlling the system performance against what they're paying for."

"For a voice agent company, they want to control their own model so that it can make sure the model actually responds by the required time... This sometimes is only you can do with your controlled intelligence... versus signing up for relying on your critical infrastructure with a proprietary API where they might go down any time."

@simonmo @BornsteinMatt

Video

a16z (@a16z)

vLLM runs on half a million GPUs at any given moment. Most people have never heard of it.

Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more.

00:00 Intro
01:46 When open source became critical infrastructure
08:55 Day zero model releases, and the drama behind them
14:59 What Kimi K3 actually buys you
18:56 Why open model licenses are changing
22:24 The pharmaceutical analogy for funding model training
26:16 If GPUs got 99% cheaper
29:08 Why vLLM, OpenRouter, and Ollama all started before ChatGPT
35:42 Building a company on an open source project
39:48 The inventor of RoPE removing RoPE

@simonmo @BornsteinMatt @VirtualElena

Video

— https://nitter.net/a16z/status/2085429677007736872#m