Reading the words is only part of reading a form. A consent form may say "Yes" and "No" next to two boxes. Plain OCR can extract both labels perfectly while losing which box contains the mark. If we flatten that output into a sentence and embed it, the retriever may return the right form but the answer generator has no reliable evidence for which option was chosen. Amazon Textract's selection-element documentation represents checkboxes separately with SELECTED or NOT_SELECTED status and links them to form fields. That illustrates the missing data type, not a guarantee that every mark is detected correctly.

I would inspect the original page and the structured extraction side by side. For the claimed consent field, which box was linked to which label, what was each box's selection state, and what page and coordinates support it? Do not infer "Yes" merely because that word appears nearer the field name in text reading order. If the box is faint, crossed out or scanned at poor quality, preserve an uncertain state and route it for review. The claim "consent granted" needs stronger evidence than a good keyword match.

In the index, store a typed field such as consent_choice = no, with source document ID, page, field label, box geometry, extraction confidence and document revision. Keep the page image available for an audit. Retrieval can find the form, but a validation step should read the structured selection and its provenance before forming the answer. Evaluate separately whether the system found the right form, linked the right checkbox to the label and correctly read the mark. One end-to-end accuracy number hides which stage failed.

The interviewer may say a vision model can look at the page directly. It can, and for ambiguous forms that may be a useful second read. Still require a cited region and a way to say "I cannot tell." The RAG answer found the right table. Why did it assign the limit to the wrong plan? covers table headers assigning a number to the wrong plan. This problem is different: the text labels are present, but the decisive information is a graphical selection state that plain OCR did not encode.