Twitter/X

Prince_Canuma congratulated Google DeepMind and announced Day‑0 MLX‑VLM native…

Brief

Prince_Canuma congratulated Google DeepMind and announced Day‑0 MLX‑VLM native diffusion decoding on Apple Silicon for DiffusionGemma, with a release expected ~3–4 hours after his 2026-06-10 18:02:11 UTC post (≈21:00–22:00 UTC) and source install available now. DiffusionGemma is a 26B MoE (3.8B active) diffusion text model that generates 256-token blocks with bi-directional attention, fits in ~18GB when quantized, and is released under Apache 2.0.

Source evidence

Massive congrats to @GoogleDeepMind on DiffusionGemma! 🎉

We collaborated closely with the team to Day-0 MLX-VLM — native diffusion decoding on Apple Silicon, release dropping later today (~3-4h), meanwhile you can install from source. ⚡🍎

This is genuinely different beast — instead of token-by-token, it generates 256-token blocks in parallel with bi-directional attention and iteratively self-corrects. 26B MoE, only 3.8B active, fits in 18GB when quantized.

Google Gemma (@googlegemma)

Meet DiffusionGemma!

An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.

Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇

Video

— https://nitter.net/googlegemma/status/2064741002204545467#m