Twitter/X

On 2026-07-28 Fish Audio announced S2.1 Pro and a $52M Seed raise, claiming it…

Brief

Fish Audio launched S2.1 Pro and announced a $52M Seed on 2026-07-28, claiming breakthrough TTS: 5‑second voice cloning, expressive word-level control, language switching mid-conversation, and open-weight public models for local GPU tinkering or API use. The thread frames this as a direct challenge to ElevenLabs, cites enterprise adopters (HeyGen, LiveKit, Retell, Sanas, OpenArt) and aggressive cost/promotional guarantees.

Why it matters

On 2026-07-28 Fish Audio announced S2.1 Pro and a $52M Seed raise, claiming it can clone a voice from 5 seconds of audio, is 2× faster than Cartesia and costs 1/6th of ElevenLabs, with word-level control over emotion, intonation, pacing and mid-conversation language switching without changing models.

Key details

  • The author asserts ElevenLabs' TTS dominance is threatened: Fish Audio offers open-weight/public-download models (the thread also claimed cloning from a 15‑second clip) so users can tinker with weights on their own GPU or use the API, and embed directives like [whispers] or [laughing nervously].
  • Fish Audio says S2.1 Pro runs in production at HeyGen, LiveKit, Retell, Sanas and OpenArt; promises to cut business voice-AI costs by 50% or give one year free, and is offering one month free for its first birthday via social engagement (like/retweet/comment “Fish”).
Source evidence

ElevenLabs founders must be SHAKING right now...

for years they owned TTS because nothing open-source sounded human

that’s over... Fish Audio just launched S2.1 Pro and they’re one of the few voice AI companies that also has open-weight models

> clone any voice from a 15 second clip
> type [whispers] or [laughing nervously] into the script and it obeys
> switch languages mid conversation without switching models

the model you used to rent a subscription for is now has a public download model

tinker with the weights on your own gpu or use the api

Fish Audio (@FishAudio)

Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.

>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc

We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.

If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: s.fish.audio/tmapke

To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.

Video

— https://nitter.net/FishAudio/status/2082152596739862853#m