Xai (via Future Tools)

Introducing Grok Voice Think Fast 2.0

Brief

Grok Voice Think Fast 2.0, announced July 29, 2026, is a next-generation speech-to-speech model that emphasizes parallel reasoning, lower latency, and improved transcription accuracy. In Artificial Analysis benchmarks it posts an AA quality index of 82.9%, speech-reasoning BigBench audio 97.2%, Conversational Dynamics 95.1%, and a τ-voice agentic score of 56.5%, while reducing median time-to-first-audio to 0.70s. Evaluations across thousands of short phrases in 24 languages report 1.5–2.0× transcription accuracy gains over Deepgram Nova 3 and ElevenLabs Scribe v2 (and ~1.4× vs Think Fast 1.0), with the advantage growing to ~10× in noisy or telephony-compressed audio. The model uses reinforcement learning to shorten turns and ask focused questions, and is more token‑efficient for reasoning (P50 = 0.4×), enabling tool calls to often finish before the agent’s first sentence ends. Automatic rollout to 'grok-voice-latest' is scheduled for Aug 5, 2026; price stated at $0.08/min.

Why it matters

Grok Voice Think Fast 2.0 (announced Jul 29, 2026) achieves an AA Speech-to-Speech Quality Index of 82.9% vs 75.7% for 1.0 and 79.1% for GPT-Realtime-2.1 (source: Artificial Analysis).

Key details

  • Speech reasoning and conversational benchmarks: Big Bench Audio 97.2% (vs 97.1 for 1.0), Conversational Dynamics 95.1% (vs 77.8 for 1.0), and Agentic Performance τ-voice 56.5% (vs 52.1 for 1.0).
  • Transcription claims: 1.5–2.0× accuracy improvement over Deepgram Nova 3 and ElevenLabs Scribe v2 across thousands of short phrases in 24 languages, widening to ~10× in noisy/telephony-compressed settings.
  • Latency and deployment: median Time to First Audio 0.70s (1.25s for 1.0); P50 reasoning tokens per response reduced to 0.4× (vs 1.0×), tooling typically completes before the agent finishes its first sentence. 'grok-voice-latest' will switch to 2.0 on Aug 5, 2026; pricing is $0.08/min.
Source evidence

Jul 29, 2026

Introducing our most capable speech-to-speech voice model.

Today, we're announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.

Grok Voice Think Fast 2.0 is our most intelligent voice model yet, building on its predecessor with meaningful gains in speech reasoning, conversational ability, and tool use reliability.

Benchmark Grok Voice Think Fast 2.0 (this release) Grok Voice Think Fast 1.0 GPT-Realtime-2.1 (High) Gemini 3.1 Flash (High)
Overall AA Speech-to-Speech Quality Index 82.9% 75.7% 79.1% 69.5%
Speech Reasoning Big Bench Audio 97.2% 97.1% 96.0% 96.6%
Conversational Dynamics Full Duplex Bench 95.1% 77.8% 95.7% 74.3%
Agentic Performance τ-voice Bench 56.5% 52.1% 45.7% 37.7%
Speed Time to First Audio 0.70s 1.25s 2.98s

Source: Artificial Analysis.

Grok Voice Think Fast 2.0 outperforms even dedicated, state of the art transcription models when it comes to accuracy. In our evaluation across thousands of short phrases in 24 different languages, we've demonstrated a 1.5–2.0× improvement relative to Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4× improvement relative to Grok Voice Think Fast 1.0.

The gap between Grok Voice Think Fast 2.0 and dedicated speech-to-text models widens to ~10× in noisy settings. We've focused on making Grok Voice Think Fast 2.0 perform exceptionally well in real-world settings, with substantial background noise and telephony compression.

Word Error Rate (%). Lower is better.

Grok Voice Think Fast 2.0

Grok Voice Think Fast 1.0

Deepgram Nova 3

ElevenLabs Scribe v2

Grok Voice Think Fast models have a unique characteristic: they reason through queries while speaking. Reasoning in parallel with speech makes the model substantially smarter than other speech-to-speech models with no impact on latency.

Grok Voice Think Fast 2.0 has been trained to be very efficient with reasoning tokens relative to its predecessor. In production settings, this means that tool calls are snappier, usually executing before the end of the agent's first sentence.

Relative reasoning tokens per response (P50). Lower is better.

Grok Voice Think Fast 2.0

0.4×

Grok Voice Think Fast 1.0

1.0×

We've trained Grok Voice Think Fast 2.0 to be a better conversationalist. By using extensive reinforcement learning to push the model towards patterns we see in real human conversation, we've found that broadly the model speaks in shorter sentences, asks one question at a time, and avoids fluff. From the user's perspective, conversations with the agent are simple and fluid, even though the model is often guiding the conversation through complex workflows and thinking several steps ahead behind the scenes.

Grok Voice Think Fast 2.0 is expected to result in improved performance across almost all use cases without any edits to existing prompts. In A/B testing this model on Starlink (+1 888 GO STARLINK), we've seen a significant increase in sales conversion rate and support containment rate.

On August 5, 2026, grok-voice-latest will move from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0. No action needed to upgrade. To stay on Grok Voice Think Fast 1.0, pin grok-voice-think-fast-1.0 before then.

We believe pricing should be predictable and transparent. Grok Voice Think Fast 2.0 is priced at $0.08 / min of audio.