This model is really perfectly timed to the availability of 192GB unified memory machines.
Unsloth AI (@UnslothAI)
DeepSeek V4 Flash 0731 can now be run locally! 🐳
Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM.
V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today.
Guide: unsloth.ai/docs/models/deeps…
GGUF: huggingface.co/unsloth/DeepS…
— https://nitter.net/UnslothAI/status/2083231049434435596#m