Local now means 330GB RAM/VRAM and routing policies.
A 325GB coding model at 40 tok/s sits between hobby local and frontier APIs, exactly where batchy coding-agent work wants a controllable tier.
Unsloth AI (@UnslothAI)
You can now run Kimi K2.7 Code locally! 🌘
We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted.
Run at >40 tok/s on 330GB RAM/VRAM setups.
Run full precision on 610 GB.
Guide: unsloth.ai/docs/models/kimi-…
GGUF: huggingface.co/unsloth/Kimi-…
— https://nitter.net/UnslothAI/status/2066492839450800427#m