Model and Inference Engineering · Staff
The screen shows $1.50. Why did the voice agent say $150?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would compare the display text with the exact input sent to speech synthesis. In this scenario, the action record correctly stores 150 cents and the UI formats it as $1.50. A separate voice-preparation function strips punctuation from an amount string and passes 150 dollars to TTS. The speech engine says "one hundred fifty dollars." Nothing was wrong with the action amount or the original UI. We lost the decimal while crossing an output-format boundary.
Text-to-speech systems can interpret a numeric string in several ways. Amazon Polly's say-as documentation explains controls for how numbers and other special text are spoken. That documentation supports the need for explicit interpretation, not a claim that a specific provider turns $1.50 into $150 on its own. In this case the damaging transformation is in our preprocessing. Avoid making the assistant's free-form text the canonical source of a payment amount.
Carry typed money through the action and confirmation path: currency code and integer minor units, with a deliberate rendering rule for screen and speech. Generate the spoken phrase from that typed value, such as "one dollar and fifty cents," rather than stripping symbols from a UI string. For a high-value action, bind both display and speech confirmation to the same action ID and amount. A dry-run test can compare action payload, UI text and TTS input exactly. Listening or speech-recognition checks on produced audio can catch pronunciation errors, but automated ASR is not an authoritative monetary parser either.
The interviewer might ask whether the user hearing the wrong amount should cancel the action. Yes, the system should stop and correct the confirmation before treating it as consent. A later audit needs to distinguish what the screen showed, what audio was actually delivered and what transaction was submitted. The caller said fifteen. Why did the voice agent send fifty? is about inbound speech being misheard before an action. This is outbound confirmation diverging from a correct typed amount after the action was prepared.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →