Data and Knowledge Systems · Staff
The PDF parser extracted every sentence. Why did the summary join two unrelated paragraphs?
The question
Interview question
A research PDF uses two columns. Text extraction finds all the words, yet the RAG answer says a limitation in the right column qualifies a result in the left column. What happened before retrieval?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
PDF text is often a set of positioned drawing operations, not a stored essay with a reliable reading sequence. An extractor may emit lines in content-stream order or sort them by vertical position. For two columns, sorting every line top to bottom can alternate left and right, creating a paragraph that no reader would see. PyMuPDF's own FAQ notes that extraction may not follow reading order and that column boundaries can need explicit detection.
I would look at the rendered page next to the extracted blocks, with each block's bounding box and order. Locate the exact two sentences the answer combined. If they came from separate columns, recover a reading order within detected column regions, then handle full-width title, figures, footnotes and tables as separate layout cases. A simple sort=True may help some pages, but geometric top-to-bottom sorting alone is not proof for a mixed layout. Keep page and block coordinates through chunking so a citation can point back to what the reader sees.
The decisive test is a page where left column paragraph A continues below its first line while right column paragraph B begins at the same height. The parser should produce A before B, and retrieval chunks should not stitch their neighboring lines. Compare the answer and citation against the rendered source, not just whether every word appears somewhere in the text dump.
For a clean one-column PDF, a lighter parser may be enough. For papers with equations, figures and side notes, layout reconstruction can need specialized extraction or visual review. OCR read every word on the page. Why did RAG combine two different clauses? concerns OCR merging clauses even though their words were read. This case is specifically a two-column reading order failure, often present even when the PDF has a perfect embedded text layer.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →