my favorite detail: the demo video was made by kimi K3 itself, running on the same SGLang deployment being announced. the model demonstrating its own serving infrastructure is a nice recursive touch. on the engineering: 2.8T MoE with a novel KDA architecture at 423 tok/s on day 0 required co-designing kernels, caching, and speculation around an attention mechanism that didnt exist in any serving framework a month ago. eleven cloud partners serving it from launch. this is what a mature open source ecosystem looks like - model, infra, and cloud providers all shipping in lockstep @lmsysorg @Kimi_Moonshot @radixark
LMSYS Org (@lmsysorg)
SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles @radixark!
How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. We've passed Kimi Vendor Verifier and are ready for production!
Thanks to @KimiMoonshot, @nvidia, @AMD, @KVCacheAI, @modal, and @baseten for building this with us, and to @googlecloud, @nebiustf, @fal, @digitalocean, @runpod, @DeepInfra and @gmi_cloud for serving K3 on SGLang.
Blog, cookbook, benchmarks in the comments.
P.S. This demo video? Kimi K3 made it itself. Play the game 👇
Video
— https://nitter.net/lmsysorg/status/2081764780621337075#m