Twitter/X

atomic_chat_hq released GGUF quantized Ling‑3.0‑flash weights on Hugging Face…

Brief

Ling‑3.0‑flash quantized weights (GGUF) are published on Hugging Face from BF16 to 1‑bit plus NVFP4. For 128GB systems tested on DGX Spark, AD‑Q5KM is the best fit (97.5% token match, 31% less drift vs llama.cpp). Ant Ling also released INT4 and FP4 (MXFP4) variants: FP4 defaults to W4A16, W4A8 targets higher throughput, and the team will keep updating inference and weights.

Why it matters

atomic_chat_hq released GGUF quantized Ling‑3.0‑flash weights on Hugging Face covering lossless BF16 down to 1‑bit quant and NVFP4.

Key details

  • AD‑Q5_K_M is recommended for 128GB hardware (tested on NVIDIA DGX Spark); it matches the original model's token choice 97.5% of the time and shows 31% less drift than the llama.cpp default quant of the same size.
  • Ant Ling released INT4 and FP4 (MXFP4) Ling‑3.0‑flash variants that run end‑to‑end on a single NVIDIA DGX Spark via a Spark‑adapted SGLang path; FP4 uses W4A16 as the stable default and W4A8 for higher throughput, with inference code and weights to be iteratively improved.
Source evidence

Run Ling 3.0 Flash locally 🌀

We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4!

AD-Q5KM is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp default quant of the same size.

Ant Ling (@AntLingAGI)

🚀 Today, we’re releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.

Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.

For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.

The efficiency and accuracy of the quantized models are still evolving. We will continue updating the inference implementation and quantized weights over time.

— https://nitter.net/AntLingAGI/status/2085024077434196211#m