Twitter/X

DiffusionGemma (Google Gemma) generates 256 tokens in parallel using…

Brief

Google’s DiffusionGemma is an experimental open-source text-generation model (Apache 2.0) that uses Bi-Directional Attention to generate 256 tokens in parallel, claiming roughly 4× speedups. Reported benchmarks show 1000+ tok/s on an H100 and 700+ tok/s on an RTX 5090; it’s a 26B Mixture-of-Experts with 3.8B active weights and fits in 18GB VRAM.

Why it matters

DiffusionGemma (Google Gemma) generates 256 tokens in parallel using Bi-Directional Attention, claimed to be ~4× faster than sequential token generation.

Key details

  • Reported throughput: 1000+ tokens/sec on a single NVIDIA H100 and 700+ tokens/sec on an RTX 5090; model is a 26B MoE with only 3.8B active parameters and fits in 18GB VRAM.
  • DiffusionGemma is released as an experimental, open model under the Apache 2.0 license and aims to produce entire blocks of text simultaneously instead of token-by-token.
Source evidence

Google’s new DiffusionGemma model is 4x faster as it can generate 256 tokens in parallel with its Bi-Directional Attention.

> 1000+ tok/s on a single h100
> 700+ tok/s on an rtx 5090
> 26B MoE, only 3.8B active, fits in 18GB vram https://nitter.net/t.co/RutyUIYxJe

Video

Google Gemma (@googlegemma)

Meet DiffusionGemma!

An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.

Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇

Video

— https://nitter.net/googlegemma/status/2064741002204545467#m