Twitter/X

Andrew Feldman asserts that inference speed equals productivity and that "fast…

Brief

Andrew Feldman claims Cerebras delivers world-leading inference speed, arguing "fast tokens are the most valuable tokens." He says Google DeepMind's Gemma 4 achieves 1,500+ tokens/second on Cerebras — over 10× faster than GPUs and 15× faster than Claude Haiku 4.5 — and notes a private preview begins today with GA later this month for multimodal app builders.

Why it matters

Andrew Feldman asserts that inference speed equals productivity and that "fast tokens are the most valuable tokens," positioning Cerebras as the maker of the fastest tokens.

Key details

  • @GoogleDeepMind's Gemma 4 runs at 1,500+ tokens/second on Cerebras hardware — claimed to be more than an order of magnitude faster than on GPUs and 15× faster than Claude Haiku 4.5.
  • Cerebras availability: private preview starts today with general availability scheduled later this month; promoted for building "blazing fast multimodal apps."
Source evidence

Fast tokens are the most valuable tokens.
Because inference speed is productivity.
And @cerebras makes the fastest tokens in the world.

@GoogleDeepMind's Gemma 4 runs at 1,500+ tokens per second on Cerebras — more than an order of magnitude faster than on GPUs and 15x faster than Claude Haiku 4.5.

Private preview today. GA later this month.
Build blazing fast multimodal apps.

Video