Twitter/X

DeepSeek released accelerated model builds; the standout is Gemma4-12B (12B…

Brief

DeepSeek published accelerated versions of several models; the author highlights Gemma4-12B (12 billion parameters) — likely with vision — as the most interesting and possibly the strongest local 12B-class model. Qwen 3.5 is absent from the accelerated lineup, reportedly because DS/DSpark omits linear-attention support.

Why it matters

DeepSeek released accelerated model builds; the standout is Gemma4-12B (12B parameters) which the author presumes includes vision and may be the best local model in its weight class by some margin

Key details

  • Qwen 3.5 was not included in DeepSeek/DSpark's accelerated set, apparently because DS[park] does not implement linear attention
Source evidence

Good guy DeepSeek gives us accelerated models
The most interesting one here is Gemma4-12B, I presume vision included. Might be the best local model in its weight class now, by some margin
Qwen 3.5 not included because DS[park] doesn't do linear attention I guess

Florian Brand (@xeophon)

🫪

— https://nitter.net/xeophon/status/2071215491641741361#m