Twitter/X

Author (@shiri_shh) is on "day 1" of self-hosting Kimi K3, citing the official…

Brief

The post shows someone beginning to self-host Kimi K3 because the vendor API seemed costly. Kimi.ai publicly released Kimi K3 — a 2.8 trillion-parameter MoE with vision capabilities and a 1M-token context, claiming 2.5× intelligence per compute — plus weights, a technical report, and open-source attention/MoE/infra tools to run it at scale.

Why it matters

Author (@shiri_shh) is on "day 1" of self-hosting Kimi K3, citing the official API as "felt expensive" as the motivation.

Key details

  • Kimi.ai released Kimi K3: a 2.8T Mixture-of-Experts model with native visual understanding, a 1M-token context window, and a claimed 2.5x intelligence per unit of compute; they also published model weights (Hugging Face), a technical report (GitHub), and supporting stack components (attention kernels, MoE communication library, agent infra).
Source evidence

POV: day 1 of self-hosting Kimi K3 because the API "felt expensive"

Video

Kimi.ai (@Kimi_Moonshot)

Releasing the model weights and technical report of Kimi K3.

Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.

New model architecture: 2.5x the intelligence per unit of compute, not just more params.

Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.

Model weights: huggingface.co/moonshotai/Ki…
Tech report: github.com/MoonshotAI/Kimi-K…
Tech blog: kimi.com/blog/kimi-k3

— https://nitter.net/Kimi_Moonshot/status/2081760186235289764#m