Twitter/X

Kimi K3 is a 2.8T-parameter Mixture-of-Experts (MoE) model with native visual…

Brief

Kimi K3 is a 2.8-trillion-parameter MoE model with built-in vision and a 1M-token context window; Moonshot says a new architecture yields 2.5× intelligence per compute. They published model weights on Hugging Face, a technical report on GitHub, a tech blog, and released attention kernels, an MoE communication library, and infrastructure for running agents at scale.

Why it matters

Kimi K3 is a 2.8T-parameter Mixture-of-Experts (MoE) model with native visual understanding and a 1,000,000-token context window.

Key details

  • Moonshot claims a new architecture that delivers 2.5× intelligence per unit of compute and released model weights (Hugging Face), a technical report (GitHub), and a tech blog; they also open-sourced high-performance attention kernels, an MoE communication library, and agent-environment infrastructure.
Source evidence

Releasing the model weights and technical report of Kimi K3.

Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.

New model architecture: 2.5x the intelligence per unit of compute, not just more params.

Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.

Model weights: huggingface.co/moonshotai/Ki…
Tech report: github.com/MoonshotAI/Kimi-K…
Tech blog: kimi.com/blog/kimi-k3