The label is missing its time and evidence. “30 days” can be correct for a historical April question, wrong for a question asking the current policy, or inapplicable to a purchase whose governing policy depends on purchase date. I would not change the label blindly to 14 either. First define what the user asked, the policy's effective dates, the source revision available to the system, and the business rule for which policy applies. Document publication time and policy effective time are different things.

Each case should carry the question, intended as-of or transaction time, expected decision and rationale, source ID and revision, effective interval, allowed evidence, permission context, and the version of the adjudication. Keep a historical replay suite that pins the April corpus, tool behavior, and expected April answer. Keep a current-policy suite whose cases and labels are adjudicated against the current authorized corpus. A case can appear in both with different setup and expected answer, but the metric must say which world it is scoring. W3C's provenance model is a useful vocabulary for connecting a decision to the source and activity that produced it. The exact schema here is an application choice.

When a document changes, do not quietly overwrite all labels. Detect affected cases through source lineage and claim dependencies, queue them for review, and classify what changed: answer, evidence span, applicability, or access. Some labels remain valid while their citation span moves. Others need a new expected answer, an explicit abstention, or retirement because the question has no stable meaning. Re-adjudicate a sample of unaffected cases too, since dependency links can be incomplete.

The evaluation report should separate model change from world change. Compare two models on the same pinned world to see whether the model improved. Run the current suite to see whether the shipped system answers today's users correctly. If the index is older than the policy source, report freshness lag and the number of requests exposed to it. A single time series mixing revised labels, new documents, and changed models is not a clean regression signal.

Now the interviewer removes the April page from the source system. The historical test can retain a licensed, access-controlled snapshot if retention rules allow it. If the document must be deleted, preserve only permitted metadata and retire any replay that would require the content. Do not resurrect a deleted policy in a production index merely to keep an eval reproducible.

Finally suppose an employee who could read the April page loses access in June. A historical correctness test can still ask what the system would have answered under the April permission snapshot, if that snapshot may lawfully be retained. A current user-facing test must use June permissions and may need to abstain or cite a different authorized source. A gold answer is not a permission grant.