Twitter/X

Solaria-3 is Gladia_io's new ASR model targeted at real-world production audio…

Brief

Solaria-3 is Gladia_io's new ASR model targeted at real-world production audio problems (noise, accents, overlaps) and core European languages; reported 2026-06-12 benchmarks put it at #1 with 9.6% WER on real English customer calls, 6.4% on Earnings22, and 33.9% on conversational telephone speech, and demos claim dramatic transcription and Audio-to-LLM improvements.

Source evidence

🚨 Most speech-to-text models look great on clean benchmarks.

Then they hit real customer calls.

Once you introduce background noise, thick accents, fast speech, or overlapping speakers, the output often becomes unusable.

That is exactly why @Gladia_io just built Solaria-3 🧵↓

It's a completely new ASR model designed for the messy production audio that actually matters to businesses.

While Solaria-1 remains Gladia's universal model for 100+ languages and live streaming,

Solaria-3 is their new precision tool for core European languages.

.. and this week's new benchmark results are absolutely WILD:

🏆 9.6% WER on real English customer calls (#1 in the field)
🏆 6.4% WER on the Earnings22 business dataset (#1)
🏆 33.9% WER on conversational telephone speech (#1)

If you are building meeting assistants, voice agents, or contact center workflows, transcription accuracy is everything.

I got a look at a few demos from the @Gladia_io team, and honestly, the results are mind-blowing.

Solaria-3 picked everything up correctly:

  • the tricky names
  • the numbers
    even niche industry jargon that usually trip up other models!

The result?

Flawless transcripts.

Massive improvements in downstream LLM performance.

Voice workflows you can actually trust.

While I haven't had the chance to put it through my own stress test just yet, Solaria-3 genuinely feels like a game-changer for Audio-to-LLM pipelines 👀↓

Video