Twitter/X

On 2026-07-27 Agent Arena ranked open-weight Kimi K3 (Max) #1 with +9.75%…

Brief

Kimi K3, a 2.8T MoE model with native visual understanding and a 1M-token context window, topped Agent Arena on 2026-07-27 with +9.75% net improvement across five signals (vs. GLM-5.2 Max at +7.12%). Moonshot published weights, a technical report, and supporting libraries; Agent Arena uses causal tracing on millions of long-horizon agentic tasks.

Why it matters

On 2026-07-27 Agent Arena ranked open-weight Kimi K3 (Max) #1 with +9.75% net-improvement across five signals, ahead of GLM-5.2 (Max) at +7.12%.

Key details

  • Kimi K3 is a 2.8T MoE model with native visual understanding and a 1M-token context window; Moonshot released model weights, a technical report, high-performance attention kernels, an MoE communication library, and agent infrastructure.
Source evidence

Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1 spot across 5 signals (see below).

Kimi K3 (Max) is also now #1 in open-weight in the Frontend Code (1682 pts) and Text (1485 pts) Arenas.

Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model.

Congrats to the @Kimi_Moonshot team for their contribution to the open ecosystem.

Kimi.ai (@Kimi_Moonshot)

Releasing the model weights and technical report of Kimi K3.

Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.

New model architecture: 2.5x the intelligence per unit of compute, not just more params.

Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.

Model weights: huggingface.co/moonshotai/Ki…
Tech report: github.com/MoonshotAI/Kimi-K…
Tech blog: kimi.com/blog/kimi-k3

— https://nitter.net/Kimi_Moonshot/status/2081760186235289764#m