Evaluation and Quality · Staff
The chart says growth, but the axes were cropped. What can the assistant claim?
The question
Interview question
An executive uploads a screenshot of a usage chart and asks, “Did adoption accelerate after the launch?” The model says yes because the line rises steeply. The crop removes the y-axis, x-axis labels, and legend. Some plotted points might be weekly totals, while others might be cumulative. Decide what the assistant can say, how to obtain a stronger answer, and how you would evaluate this capability. The original data table is unavailable for one customer.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would not claim acceleration from that screenshot. A line that rises tells us very little without the measured variable, units, time intervals, and series identity. If it is cumulative adoption, the line can rise while new adoption per week slows. Acceleration means the rate of change increased over comparable intervals, not merely that the value increased. A cropped y-axis can change the visual impression of magnitude, and irregular x-axis spacing can make a steady change look sudden. The missing legend might mean the steep line is a different customer cohort altogether.
The assistant can say what is actually visible: “One plotted line rises from left to right, but the crop hides the axes and legend, so I cannot tell whether adoption accelerated after launch.” Then ask for the full chart or underlying values, and for the launch date if it is not marked. If those are supplied, read the variable and units, compare equal time windows before and after, compute the appropriate rate or slope, and check whether a new cohort definition or instrumentation changed at launch. A correlation in the chart still would not prove the launch caused the change. For an executive decision, give the observed trend and the causal limit separately.
If the original table exists, use it as the arithmetic source and the chart as a visual cross-check. Chart-to-table extraction can help when pixels are all we have, but the conversion has its own errors. ChartQA studies questions needing both visual and logical chart reasoning, and DePlot separates plot extraction from later language reasoning. Neither paper says a model may invent cropped axes. A traceable answer should tie each extracted series and value to visible marks or to an authorized source table, with uncertainty where the resolution is too low.
For the customer with no surviving table, we cannot silently promote pixel estimates to exact business numbers. Ask for a full-resolution uncropped chart, including legend, units, time axis, plot boundaries and any note about normalization. If that too is unavailable, limit the answer to visible geometry and state what is unknown. Sometimes a qualitative statement survives: if two labeled points are visible with dates and the same scale, we may say the plotted value appears higher at the later point. But “accelerated” still needs at least comparable intervals and a defined measure. A good answer may be a careful refusal of the stronger claim with a useful request for what would settle it.
Evaluation should not use only ordinary chart questions where the answer is printed plainly. Build paired cases from the same underlying chart: full image, crop hiding the axis, crop hiding the legend, swapped legend, linear versus log scale, cumulative versus per-period plot, and different launch markers. Score claim-level support and the decision to ask for missing evidence. If a model answers all pairs with the same confident sentence, it is using the picture's shape as a shortcut. Human graders need the original data and the exact crop the model saw, otherwise they may judge a plausible answer correct without noticing that the model lacked the evidence.
There is a product fix too. If our product creates the chart, publish its data table or a meaningful long description with the image. W3C's guidance on complex images shows how charts can carry text and tabular descriptions. That helps accessibility and gives the assistant a less ambiguous source. For a third-party screenshot, we only have what the user uploaded. We should not let the model's fluency fill in the missing axes.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or return to the full Interview Prep index.
Browse this area →