Model and Inference Engineering · Staff
The PDF page is high resolution. Why did the vision model miss its small footnote?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The original file resolution tells us little about what the model actually saw. The page might be rasterized at one size, resized again by an image processor, then divided into patches or crops under an image-token budget. Eight-point text at the bottom can become too few pixels to read after a full page is squeezed down. Hugging Face's image processor documentation describes resizing as part of model-specific preprocessing. The exact limits differ by vision model and service. I would inspect the processed image or crop layout rather than guess from the PDF's DPI or the provider's “high detail” label.
Make a small probe page where the footnote changes the answer, such as “All refunds are allowed” followed by a footnote excluding a specific region. Send the full page and then a lossless crop of that footnote at readable scale. If the crop works and the whole page fails, measure the text height after every transformation and confirm whether the footnote region was ever sent. If both fail, inspect rasterization, rotation, contrast, OCR text extraction and the model's actual visual capability. A correct answer from a crop is not proof that the normal production path had access to it.
A practical pipeline can combine native PDF text extraction, OCR for scanned sections and targeted crops for small tables or notes. Keep coordinates back to the source page and attach the crop or text span used as evidence. When the footnote changes eligibility, the assistant should not answer from a headline it read confidently while the exception remains unreadable. It can ask for a better scan or mark the answer incomplete. More image patches cost latency and tokens, so target the risk-bearing regions rather than indiscriminately sending every page at maximum resolution.
The easiest release test is not “does the model read a clean paragraph.” It is whether it notices the small exception that changes a decision. If it cannot, the system should carry that uncertainty forward rather than turn an unreadable footnote into a confident answer.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →