Twitter/X

Sonic 3.5 is Cartesia’s new text-to-speech model, released as the latest…

Brief

Cartesia announced Sonic 3.5, a text-to-speech model focused on faster, more natural output for practical voice-agent use. The developer tested it specifically on multiple languages, numbers, and words that AI voices typically mishandle to evaluate robustness and accuracy; a video demonstration of those tests was provided.

Why it matters

Sonic 3.5 is Cartesia’s new text-to-speech model, released as the latest iteration in their Sonic line.

Key details

  • The model is designed to produce faster, more natural-sounding voices aimed at real voice-agent tasks.
  • Cartesia tested Sonic 3.5 on multilingual text, numeric sequences, and commonly mispronounced or error-prone words; a demonstration video accompanies the post.
Source evidence

Sonic 3.5 is Cartesia’s latest text-to-speech model.

It is built to make AI voices sound faster, more natural, and more useful across real voice-agent tasks.

So we tested it with languages, numbers, and the kind of words AI voices usually get wrong.

Video