Twitter/X

DramaBox is presented as the first "directable" speech engine and accepts stage…

Brief

DramaBox is a directable speech engine that accepts theatrical stage directions (sighs, pauses, whispers) placed outside quoted text to shape delivery. It generates 48 kHz stereo, studio-quality audio, enables zero-shot casting by description or cloning from 10 seconds of audio, embeds PerTh watermarks that survive compression/editing, and is open-source on Hugging Face.

Why it matters

DramaBox is presented as the first "directable" speech engine and accepts stage directions placed outside quotation marks (e.g., sighs, pauses, whispers) to control performance.

Key details

  • Audio output is "cinematic": 48 kHz stereo studio-quality sound and supports zero-shot casting by describing age, accent, and mood or by cloning any voice from just 10 seconds of audio.
  • Every output is PerTh watermarked by default (claimed to survive compression and editing) and the project is open source on Hugging Face (huggingface.co/ResembleAI/Dr…).
Source evidence

DramaBox is the first directable speech engine. It delivers:

>> direction: use stage directions outside the quotes (sighs, pauses, whispers) to direct the scene.

>> cinematic quality: 48 kHz stereo audio that sounds like a studio recording.

>> zero-shot casting: describe the age, accent, and mood—or clone any voice from just 10 seconds of audio.

>> verifiable Art: Every output is PerTh watermarked by default. It survives compression and editing, proving it’s your creation.

>> open source: available today on @huggingface - huggingface.co/ResembleAI/Dr…