Twitter/X

Cartesia achieved #1 rankings for both speaking and listening in the same release…

Brief

Cartesia announced a "full stack" voice AI milestone on 2026-06-15: Sonic-3.5 (streaming text-to-speech) and Ink-2 (streaming speech-to-text) are claimed as the #1 streaming models for speaking and listening, respectively. Karan Goel framed the wins as driven by new architectures that enable higher speed and quality, and positioned Cartesia as the only provider holding top rankings in both TTS and STT in the same release window.

Why it matters

Cartesia achieved #1 rankings for both speaking and listening in the same release window (announcement dated 2026-06-15).

Key details

  • Karan Goel (@krandiash) announced Sonic-3.5 (streaming text-to-speech) and Ink-2 (streaming speech-to-text) as the #1 streaming models available for voice agents, citing new architectures that boost speed and quality.
  • Cartesia claims to be the only provider with top-ranked models for both TTS and STT simultaneously.
Source evidence

Cartesia just did something no voice AI company has done before:
hit #1 for speaking and #1 for listening in the same release window.

Voice AI just got its full stack moment.

Karan Goel (@krandiash)

We released Sonic-3.5 and Ink-2, the #1 streaming models for text to speech and speech to text you can use in your voice agents today.

New architectures enable new frontiers for speed and quality.

We're now the only provider to have #1 models for both speaking and listening.

Video

— https://nitter.net/krandiash/status/2066559212533190917#m