I would look at the crop and the box in both coordinate spaces before grading the model. Reading the correct number and pointing elsewhere is possible even with a perfect OCR result if a box measured in the resized crop is drawn on the full original page as though its origin and scale were unchanged.

Take a concrete case. The source image is 2,000 by 1,200 pixels. We crop from source origin (600, 300) with size 800 × 400, then resize that crop to 400 × 200 for OCR. The OCR box center (100, 50) in the resized crop maps back to (800, 400) in the source image: first multiply each crop coordinate by two, then add the crop's source origin. Drawing (100, 50) directly on the source page produces a plausible looking but entirely wrong citation. A full box needs the same transform applied to both corners. Rotation, deskew, padding and aspect-ratio preserving resize add steps. Store the actual transform sequence or an invertible mapping rather than guessing it from final dimensions.

A crop-local OCR box maps to a different location on the source page
A citation box must be mapped through the crop and resize before it can point to evidence on the source page.

The representation needs an identity as well as geometry. Attach a source image revision and page ID to the crop, OCR spans and citation. If the source was rerendered at another DPI or replaced after OCR, a mathematically correct transform for the old raster still points to the wrong pixels on the new one. Torchvision's transforms treat boxes and images together for this reason. An OCR vendor's box might already be in source-image coordinates, crop coordinates or normalized coordinates, so check the API contract before applying any mapping twice.

I would log one example end to end: source revision and orientation, crop rectangle, resize dimensions, OCR box, model-visible image, selected text span and final source box. Overlay the mapped box on the exact source raster the user will see. Test points at corners and edges, since a center-only check can miss mirror or rotation mistakes. Include a crop that was padded and one that was rotated. If the transformed box now surrounds the right number, the coordinate pipeline was wrong. If it surrounds a different value and the model still cited it, the evidence selection is wrong. Both failures can occur together.

If we cannot maintain the mapping, I would show a page-level citation or say the exact region could not be located. A tight red rectangle feels precise to a learner, which makes a wrong one more misleading than a less specific honest citation. The answer text and the visual evidence have to refer to the same source revision and the same bytes.