Run Ling 3.0 Flash locally 🌀
We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4!
AD-Q5KM is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp default quant of the same size.
Ant Ling (@AntLingAGI)
🚀 Today, we’re releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.
Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.
For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.
The efficiency and accuracy of the quantized models are still evolving. We will continue updating the inference implementation and quantized weights over time.
— https://nitter.net/AntLingAGI/status/2085024077434196211#m