Twitter/X

Fish Audio serves 8 million users, its GitHub repo has 30K+ stars, and the…

Brief

Fish Audio used an open-models distribution strategy and heavy infra optimization to scale: they rebuilt inference on custom FP8 kernels to exceed 8,000 tokens/sec on one H200, enabling economically viable voice AI. With 8M users, 30K+ GitHub stars and a $52M Seed, they launched S2.1 Pro (83+ languages, <90ms TTF audio, 5s cloning, word-level expressive control).

Why it matters

Fish Audio serves 8 million users, its GitHub repo has 30K+ stars, and the company announced a $52M Seed raise while publicly launching S2.1 Pro.

Key details

  • They open-sourced their core models and rebuilt their serving stack on custom FP8 kernels, claiming >8,000 tokens/sec on a single NVIDIA H200 to make character-billed voice inference economically viable.
  • S2.1 Pro supports 83+ languages, delivers <90ms time-to-first-audio, can clone a voice from 5 seconds of audio, offers plain-text controls like [whisper] and [laugh], claims 2x speed vs Cartesia and ~1/6th the cost of Eleven Labs, and is used in production by HeyGen, LiveKit, Retell, Sanas, and OpenArt.
Source evidence

Fish Audio is used by 8 million people and has 30K+ stars on its GitHub repo.

If you ask them how they did it, the answer is pretty simple: by making their core models open.

It’s a similar playbook to Llama, DeepSeek, and Qwen: release open models first, then build a commercial ecosystem around them.

Voice is a tougher category for that strategy because speech gets billed by the character, so inference costs become a major part of the business model. Fish Audio rebuilt its serving stack on custom FP8 kernels, pushed past 8,000 tokens per second on a single H200, and the economics started to work.

A year later came S2.1 Pro: 83+ languages in one model, under 90ms time-to-first-audio, and plain-text control over delivery with prompts like [whisper] and [laugh]. Developers can use the hosted API or choose Fish Audio’s open models if they want to self-host. According to Fish Audio, its pricing is about 1/6th of its comparable providers.

Open source is a distribution strategy that only works when your unit economics can carry it.

Fish Audio (@FishAudio)

Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.

>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc

We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.

If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: s.fish.audio/tmapke

To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.

Video

— https://nitter.net/FishAudio/status/2082152596739862853#m