When Low Word Error Rate Still Breaks a Voice Product
One of the most dangerous speech-to-text failures is also one of the easiest to miss. The transcript reads cleanly. Then you inspect the value the workflow actually needs: DHCP became DHP, an IP address lost an octet, or a spelled command was normalized into an ordinary word.
Unlike a transcription mistake that messes up "thinking" and "tinkering," mistakes of this critical value cannot be recovered later by a large language model in the voice agent system. If the phone number is transcribed incorrectly, the large language model would have no way of recovering the correct value.
With that in mind, word error rate has earned its place as the standard measure of automatic speech recognition. In the standard NIST formulation, WER counts substitutions, deletions, and insertions relative to the reference transcript. It is simple, reproducible, and useful for comparing broad transcript fidelity or catching regressions over time.
But WER assigns roughly the same unit cost to errors with radically different consequences. Deleting a filler word and changing one digit in an account number can each contribute a single error. One makes a sentence slightly less faithful. The other sends a workflow to the wrong direction. Averaging hides that difference, especially when a long, fluent transcript surrounds one incorrect value.
That is not a flaw in WER. It is a flaw in using WER as the only release or procurement metric. WER scores the transcript. Production teams need to measure the outcome.
Semantic similarity does not close this gap. AB-410 and AB-401 look nearly identical, but they can point to different customers. An ASR system can normalize the spoken letters "P-I-N-G" into "ping," even when the spelling itself is the instruction. For identifiers, URLs, dates, measurements, file paths, and command-line flags, nearly right is wrong.
The answer is not rigid character-for-character matching. Harmless formatting should be normalized: 415-555-0123 and (415) 555-0123 can represent the same phone number. What matters is whether the correct canonical value remains recoverable without guessing. A missing digit, symbol, unit, or separator should fail when it changes the value.
The Ranking Changes When the Metric Changes
The difference is visible in Voice Code Bench, a public benchmark we helped build for structured production speech. It contains 300 human-recorded, scripted English clips with 1,482 audited target entities across 26 types and eight workplace domains. The scenarios and values are synthetic with no personal identifiable information.
Alongside WER, the benchmark reports critical token/entity match, or CTEM: whether the correct canonical value for each target remains recoverable from the transcript. It also reports task success rate, or TSR, a recording-level measure that requires every target entity in a clip to be recoverable. TSR here is not end-to-end business success; it measures complete critical-value recovery.
In the September 2026 update, the model with the lowest format-invariant WER scored 4.3 percent, yet achieved recording-level success on only 50.3 percent of clips. A different system had a higher WER of 8.2 percent but reached 68.7 percent recording-level success—an 18.3-point advantage.
WER and recording-level success ranked these two systems in the opposite order.
Build the Evaluation Stack Around the Workflow
What's the takeaway here? You should start with the workflow you are shipping, not a generic leaderboard. List the values that cannot be wrong: customer IDs, phone numbers, dates, amounts, addresses, URLs, medication names, part numbers, or commands. Build a compact regression set from real failure patterns and representative operating conditions while protecting customer data.
Keep WER or character error rate as the first layer. They remain useful measures of general transcription quality. Add critical-value recovery as the second layer, reported by entity type. Publish the canonicalization rules so harmless formatting differences pass while consequential changes fail.
The third layer is complete-request recovery. When an utterance contains four fields needed to complete an action, did all four survive? Averages can conceal compounding risk: a system that is strong on individual fields can still fail many requests if each request depends on several of them.
The fourth layer is end-to-end task success. Did the system identify uncertainty, ask for confirmation when necessary, call the correct tool, write the correct value, and complete the intended action? This layer applies whether the architecture exposes a transcript or moves directly from speech to an agent.
Finally, slice the results by the conditions that define production: entity type, accent, language, channel, noise, device, streaming versus batch, and latency. Track silent-error rate and clarification rate alongside success. Version the audio, normalization rules, model configuration, prompts, and evaluation date. When production exposes a new failure, add it to the regression set.
WER belongs on the dashboard. It should not be the only gate. The release question is not merely how many words were correct? It is, Did the system preserve what the next step needed and complete the job without a silent error?
A production voice system succeeds when it does the job, not when its transcript merely looks right.
Yi Zhong is co-founder and CEO of Besimple AI, a company focused on licensed human audio data and evaluation for voice artificial intelligence.