The model does not necessarily see the original pixels at their original scale. The client may compress, resize or tile the image, and the selected vision detail mode may discard the small glyphs that distinguish 8 from 0. Once those strokes are gone, a language model can produce a plausible code from surrounding context but cannot recover the missing bits by reasoning harder. OpenAI's image-input guide calls out small text and suggests enlarging it. Anthropic's vision guidance also recommends making important text legible and discusses resizing. The exact processing depends on the model and request path, so inspect what this pipeline actually sent and what the model actually received.

I would preserve the original image, transformed image and crop coordinates, plus request detail settings. Compare the code at each stage. Try a lossless crop of the dialog at enough resolution for the characters and pass both the crop and wider screenshot for context. A dedicated OCR pass can return candidate text and boxes, but it also makes mistakes. For a decision-critical code, compare OCR with a vision read, expose uncertainty and ask the user to confirm when the image is ambiguous. Repeatedly prompting the same low-resolution bytes is not independent evidence.

The crop must be tied back to its original region. If a dialog contains a code in its title and another in a log pane, a perfectly read crop of the wrong region still gives a wrong answer. Keep coordinates through rotation and resize, and evaluate exact code match on screenshots with varied display scaling, compression, dark themes and similar glyphs. Include the case where a screenshot is sharp on the user's device but the uploaded representation is not.

What if a higher-resolution path doubles image tokens and latency? Route it when fine text matters, rather than for every image. If the code triggers an operational action, do not turn a low-confidence visual guess into a tool parameter. The vision model reads the number correctly. Why is its citation on the wrong part of the image? covers coordinate mapping after cropping, while The PDF parser captured every word. Why did the answer reverse the chart's trend? covers extracting a chart's meaning from a parsed PDF. Here the failure starts at image detail and character fidelity, before the assistant has a reliable code to reason about.