Deepdub Launches Phantom Z 3.4 Conversational
Deepdub, a voice artificial intelligence company, has launched Phantom Z 3.4 Conversational, a multilingual text-to-speech model with high-fidelity 48 kHz audio, improved text normalization, and extended Hebrew support.
"Every voice model sounds impressive for two minutes in a demo. Very few survive two weeks with real customers," said Ofir Krakowski, CEO and co-founder of Deepdub, in a statement. "Deployments don't stall on the 95 percent a model gets right; they stall on the misread account number, the mangled surname, the one wrong digit on a live call. We built this model for that last few percent, because in production, the last few percent is the whole product."
Phantom Z 3.4 delivers an end-to-end p95 time-to-first-audio of 150 milliseconds in real-time mode at full-range 48 kHz audio, with cross-language voice transfer from less than three seconds of reference audio. Deepdub builds and trains its own speech models from random rather than licensing them. Deepdub covers more than fifty locales and dialects verified by local voice and language experts, inside a platform supporting more than 50 locales and dialects.