ElevenLabs releases Dubbing v2 model to enhance AI voice translation quality

1 hour ago 1



ElevenLabs just shipped what might be the most significant upgrade to AI-powered dubbing since the technology left the research lab. Dubbing v2, the company’s latest model, takes a fundamentally different approach to voice translation: instead of working from text transcripts like its predecessor, it listens to the original audio performance and conditions its output directly on what it hears. The practical result is translated audio that actually sounds like the original speaker said those words in another language. Tone, pacing, emotional delivery, all preserved. How Dubbing v2 works differently Traditional AI dubbing pipelines follow a predictable sequence: transcribe the audio to text, translate the text, then synthesize new speech from the translated script. Dubbing v2 skips the middleman. By conditioning directly on the source audio, the model captures vocal characteristics that a text transcript simply cannot encode. The model supports more than 90 languages with regional variants. It can detect and handle up to 32 speakers per audio file, making it viable for everything from solo podcast episodes to ensemble-cast productions with overlapping dialogue. One of the more quietl...

Read Entire Article