An image is worth 16x16 words.
Four image 64x64
All that glitters is not gold 🏅
Qwen3.6-27B-MTP-GGUF model 3 images plain, 1 MTP find it.
DGX Spark is the prefill king, but on decode is too slow, at least for me. And don't be fooled by prefill, in real life you don't load 128K tokens while coding, you send smaller chunks that are processed and stored on SSD cache. This leads to a good interaction overall.
Consider this a preview of the results, while I fight to get 128K context working flawlessly on M5 Max and DGX Spark with llamacpp.