Twitter/X

Unsloth AI released a 325 GB Kimi K2.7 Code build (down from a 1T original, −48%)…

Brief

Unsloth AI's Kimi K2.7 Code has been shrunk from 1T to a 325 GB Dynamic 2-bit build (important layers upcasted) that runs at over 40 tokens/sec on ~330 GB RAM/VRAM machines, with a full-precision option requiring ~610 GB. The post frames this size/speed point as a controllable middle tier for batch coding-agent use cases.

Why it matters

Unsloth AI released a 325 GB Kimi K2.7 Code build (down from a 1T original, −48%) using Dynamic 2-bit quantization with important layers upcasted.

Key details

  • The quantized model runs at >40 tokens/sec on machines with ~330 GB RAM/VRAM; full-precision requires ~610 GB.
  • Author framing (Boyuan Chen) positions a 325 GB, 40 tok/s local model as a new controllable tier between hobby local setups and frontier APIs for batchy coding-agent workloads.
Source evidence

Local now means 330GB RAM/VRAM and routing policies.

A 325GB coding model at 40 tok/s sits between hobby local and frontier APIs, exactly where batchy coding-agent work wants a controllable tier.

Unsloth AI (@UnslothAI)

You can now run Kimi K2.7 Code locally! 🌘

We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted.

Run at >40 tok/s on 330GB RAM/VRAM setups.

Run full precision on 610 GB.

Guide: unsloth.ai/docs/models/kimi-…
GGUF: huggingface.co/unsloth/Kimi-…

— https://nitter.net/UnslothAI/status/2066492839450800427#m