Google’s new DiffusionGemma model is 4x faster as it can generate 256 tokens in parallel with its Bi-Directional Attention.
> 1000+ tok/s on a single h100
> 700+ tok/s on an rtx 5090
> 26B MoE, only 3.8B active, fits in 18GB vram https://nitter.net/t.co/RutyUIYxJe
Video
Google Gemma (@googlegemma)
Meet DiffusionGemma!
An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.
Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Video
— https://nitter.net/googlegemma/status/2064741002204545467#m