Smaller DeepSeek-V4-Flash quants are now available to run from 1-bit to 4-bit. Run 3-bit on 128GB RAM, 1-bit on 96GB RAM.
Remember full precision is mostly FP4 (4-bit) so 1 to 3-bit are not that quantized. Even more variants are coming soon: huggingface.co/unsloth/DeepS…
Link
unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co