Model and Inference Engineering · Staff
The call opens in English. Why did the transcript miss the Hindi complaint later?
The question
Interview question
A support call begins with an English greeting. The transcription pipeline detects language once from that opening and locks English decoding for the full call. Ten minutes later the customer explains the real problem in Hindi, using an English product name. The transcript is fluent around the greeting but garbled around the complaint, and the assistant concludes no issue was reported. Where did the evidence go?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The pipeline made a call-level language decision from a small, unrepresentative interval. Automatic language identification is an inference with scope and uncertainty, not a fact about every second of an hour-long conversation. Google Cloud's language-recognition guidance describes choosing candidate languages, and its language support varies by model and configuration. Our hypothetical pipeline's permanent lock is our design choice. The service documentation does not claim it always locks a whole call from a greeting. The model may be very good at English speech while unable to decode the later segment under that configuration.
I would inspect the original audio at the complaint interval and the ASR request configuration used for it. Was the Hindi segment present after voice-activity detection? Did the channel or speaker change? What language hypotheses and confidence were available, and were they discarded after the first segment? Compare a multilingual or segment-level decode against the English-only path, preserving timestamps and channel IDs. Product names often need a domain vocabulary or post-ASR normalization, but do not overwrite uncertain Hindi words with a plausible English paraphrase just to make a clean transcript.
A better design depends on the traffic. If calls can switch languages, admit the expected candidate languages and re-evaluate language over segments, with hysteresis so a single borrowed word does not flip the decoder back and forth. Preserve the original audio and a confidence or uncertainty flag for decision-critical spans. A human review path may be needed when the assistant is about to close a complaint as “no issue.” Evaluate Hindi-only, English-only and mixed turns, names in Latin and Devanagari, low-quality phone audio and mid-sentence switches. Measure whether the downstream task captured the complaint and took the right next step, not only overall word error rate on English greetings.
The interviewer may suggest translating every call to English first. Translation cannot recover words that ASR already missed, and it can alter names or negation. The order is capture, identify, transcribe with an appropriate model, preserve source spans, and then translate or summarize if needed. A Hindi-English query names a product in Latin letters. Why does search miss its policy? covers a Hindi-English text query missing a policy. This is an audio acquisition and decoding boundary before search begins.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →