Evaluation and Quality · Principal
The assistant won the A/B test on Tuesday. Why did the win disappear on Friday?
The question
Interview question
A team launches a new assistant and checks a conventional 5% significance test every hour. On Tuesday a dashboard shows `p < 0.05`, so they announce a win. By Friday the estimate is near zero. Was Tuesday's math necessarily wrong?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The calculation may have been correct for a fixed sample at that instant, while the decision rule was not. If the team repeatedly looks and stops as soon as ordinary fixed-horizon p-values cross 0.05, its chance of claiming a win under no real effect is higher than the nominal level. Early estimates that happen to cross a threshold are also prone to look larger than the eventual effect. Spotify's engineering discussion of sequential testing explains why a valid repeated-look procedure needs an explicit stopping rule. This is not a reason to hide safety alarms until the test ends. Monitoring guardrails and claiming a positive treatment effect are different decisions.
I would reconstruct the experiment plan: primary outcome, randomization unit, minimum duration and sample, planned looks, analysis method, guardrails and stopping conditions. Was Tuesday the predetermined final analysis or one of many dashboard refreshes? Did the team change metrics after seeing results? Did the same user stay in the same variant across turns? A multi-turn agent evaluated by request rather than user can create contamination, which What is the unit of an online experiment for a multi-turn agent? addresses separately. Here the focus is the repeated decision to declare a win from a conventional threshold.
Choose either a precommitted fixed-horizon analysis or a group-sequential or always-valid method designed for the actual monitoring schedule. If many metrics and segments are scanned for “the winner,” account for that search too. Report the effect size and uncertainty, not only a crossing of 0.05. Include costs, latency, unsafe outcomes, reversals and task completion as guardrails. For a low-base-rate harmful event, an ordinary week-long experiment may lack power to show safety, so complement it with targeted audits. A valid statistical test cannot compensate for a metric that measures the wrong user outcome.
The interviewer may say they needed to stop Tuesday because the new assistant was causing harm. Yes, predefined harm rules can stop an experiment for safety without calling Tuesday a positive win. If the system had a genuine preplanned sequential rule and crossed its adjusted boundary, Tuesday could be a valid decision. The question is which rule governed the decision before the data arrived. You tuned against the holdout for six months. Is it still a test set? covers tuning on an old offline holdout. This is online repeated peeking at a live randomized experiment.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →