Twitter/X

Kimi K3 is a 2.8T-parameter Mixture-of-Experts (MoE) model with native visual…

Brief

Kimi K3 is a 2.8 trillion-parameter MoE model with native visual understanding and a 1,000,000-token context window; MoonshotAI claims a new architecture yielding 2.5× intelligence per unit of compute. On 2026-07-27 MoonshotAI published weights and a technical report and opened high-performance attention kernels, an MoE comms library, and agent-run infrastructure; @Dorialexander praised the report.

Why it matters

Kimi K3 is a 2.8T-parameter Mixture-of-Experts (MoE) model with native visual understanding and a 1,000,000-token context window; MoonshotAI claims the new architecture delivers 2.5× the intelligence per unit of compute.

Key details

  • On 2026-07-27 MoonshotAI published Kimi K3 model weights (Hugging Face) and a technical report (GitHub) and released supporting infrastructure: high-performance attention kernels, an MoE communication library, and tooling to run agent environments at scale; @Dorialexander called it "probably the best model report this year."
Source evidence

So probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters.

Kimi.ai (@Kimi_Moonshot)

Releasing the model weights and technical report of Kimi K3.

Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.

New model architecture: 2.5x the intelligence per unit of compute, not just more params.

Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.

Model weights: huggingface.co/moonshotai/Ki…
Tech report: github.com/MoonshotAI/Kimi-K…
Tech blog: kimi.com/blog/kimi-k3

— https://nitter.net/Kimi_Moonshot/status/2081760186235289764#m