It tells us the combined stack performed better under the measured conditions. To attribute the change, run the four combinations on the same assigned questions: old retriever with old model, new retriever with old model, old retriever with new model, and new retriever with new model. Keep corpus snapshot, permission view, prompt assembly, answer budget and grading rule fixed. Then compare outcomes at the task level, not just two aggregate dashboard numbers. The old model might benefit from better evidence. The new model might handle noisy evidence better. They may also interact, so the benefit of one component can depend on which version of the other is used.

This is a small factorial design, an established controlled-experiment pattern described in the practical guide to online experiments. Offline, replaying all four combinations on a fixed test set is often the cheapest diagnosis. If the claim concerns live user outcomes, randomize the relevant combined treatments with a stable unit such as user or case, and watch for shared caches and changing documents that contaminate comparisons. Four online arms may be too expensive or underpowered. In that case run the safe offline crossing first and make a narrower live claim about the final combined system.

I would also inspect where the new retriever changes the evidence. Compare retrieval coverage, permission correctness and source version on a paired set, then judge whether the generated claims are supported by what each model actually saw. A model can appear better merely because it received easier passages. A retriever can appear better by handing the model longer context that increases cost or hides critical exceptions. The component metrics explain mechanisms, while end-to-end task success decides whether the release helps users.

If someone asks which team's work caused the win from only the original two-arm comparison, the honest answer is that we cannot know. We can still ship the combined stack if safety, latency and cost gates pass. The narrower claim is a strength: the release improved the measured task outcome, while component credit remains unproven.