Voice AI Models Mispronounce New Drug Names

Article Featured Image

Synthio Labs, a clinical-grade voice and agentic artificial intelligence company for the pharmaceuticals industry, benchmarked how accurately text-to-speech (TTS) systems pronounce drug names and found that the leading voice AI models mispronounce up to one in three new drug names.

The company tested nine commercial TTS systems on 274 medicines in clinical sentences; ElevenLabs, Google, Deepgram, Microsoft, and OpenAI models drop sharply on recently approved drugs.

In Synthio's testing, general-purpose models passed between 63.1 percent and 80.3 percent of drug names overall. On newly approved names, every model fell: ElevenLabs' eleven_v3 dropped from 93 percent on established names to 67.1 percent; Google's Gemini TTS from 89.1 percent to 61.6 percent. Microsoft Azure passed fewer than half of generic (INN) names. One system spelled Xofluza letter by letter.

"Voice models are sold on how human they sound. A model can sound human and still mispronounce the drug name, and in pharma that is the failure that matters," said Rajashekar Vasantha, co-founder and chief technology officer of Synthio Labs, in a statement. "Sound-alike drug name confusion is a patient-safety category WHO, ISMP and FDA already track. Voice AI adds a new speaker to that problem, and nobody has measured it."

DOSE scores each system's pronunciation from zero to five against verified reference pronunciations; a score of four or higher passes. Synthio Labs' own RxPronounce model is one of the nine; it passed 91.2 percent of names overall and 87 percent of newly approved names, 10.9 points ahead of the next system.

"The pattern tracks where training data runs out," said Supreet Deshpande, co-founder and CEO, Synthio Labs. "Every model handles metformin, a legacy diabetes drug, but they struggle with the molecule approved last quarter, exactly the name a launch team or a patient most needs spoken correctly."