Twitter/X

Bland Speech v3 by @usebland ranked #1 on Design Arena’s Audio Realism Bench…

Brief

Bland Speech v3 by @usebland topped Design Arena’s Audio Realism Bench (announced 2026-08-04), which asks vetted native speakers to blind-compare TTS and human recordings across phone agents, conversational assistants, and spoken explainers. Real human samples were included to test indistinguishability; MAI-Voice-2 (@MicrosoftAI) came second and Grok TTS (@SpaceXAI) third.

Why it matters

Bland Speech v3 by @usebland ranked #1 on Design Arena’s Audio Realism Bench (announcement dated 2026-08-04).

Key details

  • Audio Realism Bench uses blind pairwise comparisons by vetted native speakers across phone agents, conversational assistants, and spoken explainers, and it mixes real human recordings into the pool to test indistinguishability.
  • MAI-Voice-2 (@MicrosoftAI) placed second and Grok TTS (@SpaceXAI) placed third; Design Arena framed these results as progress toward voice systems that can sustain genuinely human-feeling interactions.
Source evidence

Congrats to the @usebland team for the most human-sounding TTS

Design Arena (@DesignArena)

BREAKING: Bland Speech v3 by @usebland debuts as the most human-sounding TTS model on our newest evaluation: Audio Realism Bench.

Traditional speech benchmarks increasingly struggle to separate frontier systems that all sound clean, fluent, and intelligible.

Audio Realism Bench instead asks a harder question: when two voices read the same transcript, which one sounds more like a real person?

Vetted native speakers evaluate recordings blindly across phone agents, conversational assistants, and spoken explainers, judging the subtle qualities that define human-like speech.

We also place real human recordings directly into the evaluation pool, testing whether evaluators can distinguish TTS models from actual human speech.

Bland Speech v3 takes the top spot in our initial rankings, followed by MAI-Voice-2 by @MicrosoftAI and Grok TTS by @SpaceXAI, all demonstrating important steps toward voice systems capable of sustaining genuinely human-feeling interactions.

Congratulations to the @usebland team!

— https://nitter.net/DesignArena/status/2084678696502468783#m