Twitter/X

Kimi K2.7 (1T) was quantized to a 325 GB GGUF using 'Dynamic 2-bit' with…

Brief

Kimi K2.7 Code (1T) has been compressed into a 325 GB GGUF via Dynamic 2-bit quantization that upcasts critical layers, cutting model size by 48%. Unsloth AI says the quantized build achieves over 40 tokens/sec on ~330 GB RAM/VRAM setups, while a full-precision run needs about 610 GB; download and guide links are provided.

Why it matters

Kimi K2.7 (1T) was quantized to a 325 GB GGUF using 'Dynamic 2-bit' with selective upcasting of important layers, a 48% size reduction from the original

Key details

  • Unsloth reports local runtimes >40 tokens/sec on systems with ~330 GB combined RAM/VRAM; full-precision requires ~610 GB
Source evidence

You need 325GB VRAM to run Kimi K2.7 locally…

Unsloth AI (@UnslothAI)

You can now run Kimi K2.7 Code locally! 🌘

We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted.

Run at >40 tok/s on 330GB RAM/VRAM setups.

Run full precision on 610 GB.

Guide: unsloth.ai/docs/models/kimi-…
GGUF: huggingface.co/unsloth/Kimi-…

— https://nitter.net/UnslothAI/status/2066492839450800427#m