Suppose a support policy changed this month. The old gold answer says a customer cannot request a refund after 14 days. The current rule allows a 30-day window for a certain plan. The new assistant applies that exception and the frozen answer key marks it wrong. A model regression and a stale evaluation target can look identical in the score.

The eval case needs more than a question and an expected string. It needs the policy version or effective date that defines correctness, the relevant customer facts, and a grader that judges the answer against that version. OpenAI's evaluation guidance emphasizes clear grading criteria and examples of score levels. That is a design requirement here: the rubric itself changes when the business rule changes.

I would audit the apparent regressions by joining each case to its source policy and effective interval. For each disputed answer, ask whether the input specified a date. If it did not, the case might be ambiguous rather than stale. Some historical questions still require the old rule, so blindly rewriting all gold answers to the new rule would also be wrong. Maintain a frozen historical slice for longitudinal comparison and a current-policy slice for release decisions, with explicit versions on both. If a grader uses retrieved policy text, pin the corpus snapshot or record the exact sources it saw.

Regrade both old and new assistants under the same current rubric to measure improvement on today's task. Keep the old rubric result separately if you need to understand behavior drift. Do not splice old scores and new scores into one trend as if the yardstick stayed fixed. Review a sample with policy owners when the source rule is subtle.

The hard pushback is, “Can we just delete outdated cases?” Delete only if they no longer test a relevant behavior. Historical reasoning, effective-date handling and conflict resolution may still matter. Relabel and version the cases so a lower score means the assistant got the task wrong, not that it knew a rule more recent than the answer key.