ElevenLabs Launches Dubbing v2 API

Article Featured Image

ElevenLabs has made its Dubbing v2 model available through its ElevenAPI platform, allowing developers to integrate the technology directly into their products and workflows.

Dubbing v2 employs a unique direct speech-to-speech (audio-to-audio) architecture that generates speech output directly from the source speech, enabling it to replicate the speaker's voice, maintaining the original's emotional expression and performance. It also aligns the spoken dialogue naturally with video timing. The model supports 92 languages.

"Dubbing v2 fixes flat audio by conditioning directly on the original performance. It ensures that tone, emotion, and delivery are preserved across 90+ languages. Sync-aware translation logic means that starts and stops align with the original out of the box," said Eric Michaelis, a member of the marketing team for ElevenLabs API, in a blog post. "Dubbing v2 now captures regional accents with sharper locale accuracy, distinguishing variants like Castilian and Latin American Spanish. It also handles background audio, music, and effects better, including scenes with multiple speakers."

With Dubbing v2, translation, voice cloning, dubbing, and sync all run automatically through a single API. For more granular editing, companies can bring their own source transcripts and target language translations or edit the ones the model produces then regenerate only the segments they change.