Twitter/X

MOSS-TTS-Local Transformer v1.5 is available on mlx-audio (announced 2026-06-20)…

Brief

MOSS-TTS-Local Transformer v1.5 is now released on mlx-audio and open sourced. The system pairs a 2B-parameter MOSS-Audio-Tokenizer-v2 with a Qwen3-4B LLM backbone to produce native 48 kHz stereo audio, streaming with theoretical sub-100 ms time-to-first-token, zero-shot voice cloning, inline pause control, and 31-language synthesis (including EN/JP/KR) with Day0 SGLang-Omni support for real-time voice-agent and digital-human applications.

Why it matters

MOSS-TTS-Local Transformer v1.5 is available on mlx-audio (announced 2026-06-20) and is now open source.

Key details

  • Model stack: MOSS-Audio-Tokenizer-v2 (2B parameters) + Qwen3-4B backbone; outputs native 48 kHz stereo audio with streaming capable of theoretical sub-100 ms TTFT and supports zero-shot voice cloning and inline [pause] control.
  • Language and integration: 31-language synthesis including English, Japanese, Korean; SGLang-Omni Day0 support noted; targeted use cases include voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation.
Source evidence

MLX news: MOSS-TTS-Local Transformer 1.5 available on mlx-audio now! Thanks to @Prince_Canuma and @lllucas

Audio on! 🔉

Video

OpenMOSS (@Open_MOSS)

🤗 MOSS-TTS-Local Transformer v1.5 is now open source.

Built with a pure autoregressive Audio Tokenizer + LLM paradigm:

>MOSS-Audio-Tokenizer-v2, 2B params
>Qwen3-4B backbone
>Native 48 kHz stereo audio
>Streaming output with theoretical sub-100 ms TTFT
>Zero-shot voice cloning
>Inline [pause] control
>🇺🇸 🇯🇵 🇰🇷 31 language synthesis
>SGLang-Omni Day0 support 🎉 @sgl_project @lmsysorg

Designed for voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation.

👇

Video

— https://nitter.net/Open_MOSS/status/2067482316696629566#m