REDDIT_POST

Best Qwen3.5-35B-A3B GGUF for 24GB VRAM?!

Brief

Qwen3.5-35B-A3B GGUF is a 24GB-VRAM-targeted quant build (posted 2026-02-25) that deliberately uses only legacy llama.cpp quant formats (q80, q40, q41); its Q40 variant is 19.776 GiB (4.901 BPW). The poster reports competitive perplexity and suggests Vulkan/ROCm backends may run the mix faster, asks for sweep benchmarks on 7900XTX/Strix Halo and Mac setups, and links a Hugging Face release.

Source evidence

title: Best Qwen3.5-35B-A3B GGUF for 24GB VRAM?!
contenttype: article
published: 2026-02-25T00:00:00
source
url: https://www.reddit.com/r/LocalLLaMA/comments/1resggh/bestqwen3535ba3bgguffor24gb_vram/

word_count: 109

Best Qwen3.5-35B-A3B GGUF for 24GB VRAM?!

My understanding is Vulkan/ROCm tends to have faster kernels for legacy llama.cpp quant types like q80/q40/q4_1. So I made a mix using only those types!

Definitely not your grandfather's gguf mix: Q4_0 19.776 GiB (4.901 BPW)

Interestingly it has very good perplexity for the size, and may be faster than other leading quants especially on Vulkan backend?

I'd love some llama-sweep-bench results if anyone has Strix Halo, 7900XTX, etc. Also curious if it is any better for mac (or do they mostly use mlx?).

Check it out if you're interested, compatible with mainline llama.cpp/ik_llama.cpp, and the usual downstream projects as well:

https://huggingface.co/ubergarm/Qwen3.5-35B-A3B-GGUF?showfileinfo=Qwen3.5-35B-A3B-Q4_0.gguf