PDF text is often a set of positioned drawing operations, not a stored essay with a reliable reading sequence. An extractor may emit lines in content-stream order or sort them by vertical position. For two columns, sorting every line top to bottom can alternate left and right, creating a paragraph that no reader would see. PyMuPDF's own FAQ notes that extraction may not follow reading order and that column boundaries can need explicit detection.

I would look at the rendered page next to the extracted blocks, with each block's bounding box and order. Locate the exact two sentences the answer combined. If they came from separate columns, recover a reading order within detected column regions, then handle full-width title, figures, footnotes and tables as separate layout cases. A simple sort=True may help some pages, but geometric top-to-bottom sorting alone is not proof for a mixed layout. Keep page and block coordinates through chunking so a citation can point back to what the reader sees.

The decisive test is a page where left column paragraph A continues below its first line while right column paragraph B begins at the same height. The parser should produce A before B, and retrieval chunks should not stitch their neighboring lines. Compare the answer and citation against the rendered source, not just whether every word appears somewhere in the text dump.

For a clean one-column PDF, a lighter parser may be enough. For papers with equations, figures and side notes, layout reconstruction can need specialized extraction or visual review. OCR read every word on the page. Why did RAG combine two different clauses? concerns OCR merging clauses even though their words were read. This case is specifically a two-column reading order failure, often present even when the PDF has a perfect embedded text layer.