Data and Knowledge Systems · Staff
OCR read every word on the page. Why did RAG combine two different clauses?
The question
Interview question
A policy PDF has two columns. The left column describes standard refunds and the right column describes exceptions for a different product. OCR recognizes all the words correctly. The extracted text interleaves lines by their vertical position, so one chunk pairs the standard rule with an exception from the other column. Retrieval finds that chunk, and the assistant confidently cites the right page but the wrong policy. Where would you fix this?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Word recognition is only one part of document understanding. A page is a two-dimensional layout. Turning it into a one-dimensional sequence requires reading order, region boundaries and table or caption relationships. If the extraction step mixes columns, downstream embedding and generation cannot reliably reconstruct the intended clause from the already flattened text. A citation to the correct page is not enough when the claimed relationship between two sentences never existed on that page. OmniDocBench includes reading-order and layout annotations separately from text recognition. Google's Document AI layout parser guidance likewise treats headers, tables and figures as layout structure rather than just a bag of words.
I would inspect three representations side by side for the failed answer: the rendered page with bounding boxes, the parser's ordered regions, and the exact chunk fed to retrieval. Find the first stage where the exception moved next to the standard rule. Check whether the PDF already has usable text and structural tags, whether OCR was necessary, and whether a generic left-to-right, top-to-bottom sorter ignored columns. A better parser may be needed, but first check if our own chunker discarded the parser's reading order or joined regions across headings. Preserve page, block, coordinates and source version so we can point back to the physical evidence.
Rebuild chunks along semantic regions: clause with its heading and qualifiers, table with relevant row and column headers, exceptions attached to the rule they actually modify. Do not blindly split every fixed number of tokens. For a policy answer, require a citation to the specific clause region, and verify that the answer's condition and product match the source. If the layout is ambiguous, send the page image or cropped regions to a vision-capable review path and allow abstention. Vision can help, but it must also be evaluated on the same hard pages. It should not be treated as an infallible tie breaker.
The interviewer may propose fixing retrieval because it returned a mixed chunk. Reindexing is necessary after parser repair, and a reranker can reject some nonsense, but it cannot guarantee correction of corrupted source text. Build a page-level test set with multi-column policies, nested exceptions, footnotes, tables and amended versions. Score reading order, clause boundaries and answer support, not only OCR character accuracy and recall@k. The failure started before search and ended as a plausible, wrong answer.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →