Model and Inference Engineering · Principal
The caller said fifteen. Why did the voice agent send fifty?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
A caller says “transfer fifteen thousand.” ASR outputs “fifty thousand.” The rest of the transcript is excellent, the LLM extracts the amount, and the payment API validates the number. The action is still wrong. A high overall transcription score tells us little about this one critical value.
Word error rate counts substitutions, deletions and insertions across a transcript. One substitution in a long call can be a small contribution to the average while multiplying the requested amount. Spoken-language systems also extract slot values such as names and amounts from imperfect transcripts. Research on ASR slot-value error recovery treats downstream value errors as a problem to handle explicitly. It does not claim a general confidence threshold that makes payment safe.
The critical boundary is between a speech hypothesis and an authorized financial action. Preserve the original audio span and timing for the amount, the ASR alternatives if available, extraction provenance, and the normalized currency and numeric value. Check whether “fifteen” versus “fifty” was uncertain in the acoustic hypothesis, whether context or an LLM rewrote the value, and whether a previous turn stated another amount. Do not let the language model silently choose the more plausible amount from account history. Plausibility is not authorization.
For this action I would read back the amount and recipient in a form that disambiguates the number, then require an explicit confirmation tied to the exact amount, currency, payee and transaction ID. If the user corrects it, invalidate the earlier action proposal and rebuild the approval artifact. The payment service checks that the confirmed tuple matches the submitted tuple. Confirmation is not merely another free-form phrase interpreted by the same noisy path. Depending on the channel and risk, use a visual confirmation, keypad entry, authenticated app approval, or a spoken challenge designed to reduce confusion. The product needs to account for accessibility and call drop-offs, but those costs should be measured against the loss from a wrong transfer.
The interviewer might suggest using a second ASR model. Independent transcription can surface disagreement, which is useful, but correlated acoustic ambiguity can fool both models. Measure critical-slot error and false authorization per amount and language slice, not just WER. Include near-sounding numbers, pauses, currency switches, code switching and noisy calls. The voice agent heard 'approve the transfer.' Did it miss the next two words? concerns words arriving after the apparent approval. This one concerns the exact value of a completed utterance before the agent ever calls the tool.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →