Twitter/X

Kimi K3 is a 2.8T MoE model that reportedly generated the demo video itself while…

Brief

Kimi K3, a 2.8T Mixture-of-Experts model, reportedly produced the demo video itself on the same SGLang deployment. LMSYS claims day‑0 throughput of 423 tok/s on gsm8k by shipping a novel KDA architecture and heavy infra co‑design (fused kernels, DP attention, DSpark, PD disagg, KDA caching). Eleven cloud providers and multiple hardware/software partners are serving it; it passed the Kimi Vendor Verifier and has RL support ready in Miles.

Why it matters

Kimi K3 is a 2.8T MoE model that reportedly generated the demo video itself while running on the same SGLang deployment being announced.

Key details

  • SGLang achieved 423 tok/s on day 0 (measured on gsm8k) by implementing a novel KDA architecture with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching, requiring co-designed kernels, caching, and speculation for a new attention mechanism introduced within the prior month.
  • Launch partners include @Kimi_Moonshot, @nvidia, @AMD, @KVCache_AI, @modal, @baseten and eleven cloud providers (e.g., Google Cloud, NebiusTF, Fal, DigitalOcean, Runpod, DeepInfra, GMI Cloud); the build passed the Kimi Vendor Verifier and LMSYS says RL support is ready in Miles (@radixark).
Source evidence

my favorite detail: the demo video was made by kimi K3 itself, running on the same SGLang deployment being announced. the model demonstrating its own serving infrastructure is a nice recursive touch. on the engineering: 2.8T MoE with a novel KDA architecture at 423 tok/s on day 0 required co-designing kernels, caching, and speculation around an attention mechanism that didnt exist in any serving framework a month ago. eleven cloud partners serving it from launch. this is what a mature open source ecosystem looks like - model, infra, and cloud providers all shipping in lockstep @lmsysorg @Kimi_Moonshot @radixark

LMSYS Org (@lmsysorg)

SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles @radixark!

How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. We've passed Kimi Vendor Verifier and are ready for production!

Thanks to @KimiMoonshot, @nvidia, @AMD, @KVCacheAI, @modal, and @baseten for building this with us, and to @googlecloud, @nebiustf, @fal, @digitalocean, @runpod, @DeepInfra and @gmi_cloud for serving K3 on SGLang.

Blog, cookbook, benchmarks in the comments.

P.S. This demo video? Kimi K3 made it itself. Play the game 👇

Video

— https://nitter.net/lmsysorg/status/2081764780621337075#m