TMCnet News
Deepdub Launches Phantom Z 3.4 Conversational: Multilingual Text-to-Speech Built to Survive Real Customers, Not Just DemosEnterprise-grade real-time text-to-speech delivers 150ms time to first audio at full 48 kHz, with text normalization that gets account numbers, invoice totals and appointment dates right TEL AVIV, Israel, Sept. 3, 2026 /PRNewswire/ -- Deepdub, a foundational voice AI company pioneering expressive voice technologies, announced today the launch of Phantom Z 3.4 Conversational, a new multilingual text-to-speech model with high-fidelity 48 kHz audio, improved text normalization and extended Hebrew support. The model is available to all Deepdub clients now. For enterprises running voice agents, a call holds together when four things go right at once. The voice sounds like a person. The response arrives fast enough to feel like a conversation. The agent knows when to speak and when to listen. And every account number, date and amount comes out the way a customer would say it. When one of them slips, the call escalates to a human, and that is where containment and cost are decided. Phantom Z 3.4 Conversational is built for all four. "Every voice model sounds impressive for two minutes in a demo. Very few survive two weeks with real customers," said Ofir Krakowski, CEO and co-founder of Deepdub. "Deployments don't stall on the 95% a model gets right, they stall on the misread account number, the mangled surname, the one wrong digit on a live call. We built this model for that last few percent, because in production, the last few percent is the whole product." In English, the work is in text normalization, the step that turns written text into spoken words. A delivery date written 2024-12-31 is read as December thirty first, twenty twenty-four rather than as a run of digits. An invoice total written $1,240 is read as one thousand two hundred forty dollars. An appointment at 14:30 is read as two thirty. A reference written Chapter VII is read as chapter seven rather than as letters. These are the categories where Deepdub's testing puts the model ahead of the other systems it was measured against. An enterprise running more than one language gets one set of behavior to test and one contract to hold rather than two. Phantom Z 3.4 delivers an end-to-end p5 time-to-first-audio of 150 milliseconds in real-time mode at full-range 48 kHz audio, with cross-language voice transfer from under three seconds of reference audio. Deepdub builds and trains its own speech models from random rather than licensing them, which allows the company to bring a new language into production in two weeks. Deepdub covers more than fifty locales and dialects verified by local voice and language experts, inside a platform supporting more than 50 locales and dialects. "We run Deepdub in production for live, real-time phone calls, where latency and naturalness aren't nice-to-haves but the key factor in whether a caller stays on the line. 3.4 is the closest we've heard a synthetic voice come to a real person, and our callers show it: they stay longer, talk more, and engage with our agents like we've never seen before," said Adir Haziza, CTO at Voiceman. The hardest case is Hebrew, which is written without vowels, so the same letters can spell different words. The three letters of ??? are a sign read one way and a remote control read another. A model that reads one word at a time has to guess which the sentence means, and in Hebrew a wrong guess is not an accent, it is a different word that stays invisible until a customer hears it. Phantom Z 3.4 resolves this at the source. Pronunciation is decided from the whole sentence rather than word by word, and every instance of ??? in Deepdub's Hebrew test set was read correctly. Where a brand name or a plan tier has to be said a particular way, marking it in the text is enough. Deepdub ranks first for Hebrew text-to-speech on the public TTS Arena leaderboard hosted by ivrit.ai on Hugging Face. In Hebrew, national ID numbers, appointment dates and transaction amounts are expanded before speech, so a balance written as 1,240 ? is spoken in full rather than read out as digits. In blind listening tests, Phantom Z 3.4 was preferred over Deepdub's previous Hebrew model in 71 percent of decisive comparisons. "We needed something that would hold up consistently across a large volume of work, so we tested it thoroughly before deciding. What stood out was that the details came out right and the Hebrew was the most natural we'd heard," said Dor Levy, Head of Jeen Talk at Jeen AI. About Deepdub Deepdub Media Contact
SOURCE Deepdub
|
