It prevented exact ID reuse. It did not separate the underlying information. A model can learn the answer from a near-duplicate training page and appear to generalize to a held-out question about the same issue. Research on deduplicating language-model training data finds substantial near-duplicate material and train-test overlap. For a targeted assistant, the problem is sharper when one answer has many copied forms. This does not prove every high score is memorization, but it weakens a claim about unseen issues.

I would trace content lineage before choosing a split key. Group pages by canonical source, product issue, version and copied text, including pages that paraphrase rather than exactly duplicate. Use exact hashes for byte copies, normalized text hashes for formatting changes, and approximate matching such as shingles or embeddings to nominate near-duplicates for human inspection. No single similarity threshold is a perfect truth test. Two pages with shared boilerplate can still contain distinct answers, while a short critical fix can be paraphrased enough to evade a whole-document hash. Inspect decisive answer spans and their timestamps.

Then define what the eval is meant to test. If the product must answer new questions about known policy documents, some overlap between the document corpus and held-out queries is expected and even necessary. If the claim is that a trained model generalizes to previously unseen product issues, group by issue family or source lineage before splitting. Keep a separate fresh set created after the training-data cutoff, with controls for public copies that may have been crawled earlier. Report both kinds of performance rather than silently calling them the same generalization result.

For containment, compute a contamination audit on the existing eval, mark affected cases, and compare scores on uncontaminated and new issues. Do not merely delete hard cases until the score looks good. Preserve split definitions, corpus snapshots and hashes so a later training run cannot ingest the holdout through a new connector. If training data is already frozen, this may require new eval rather than pretending we can untrain a model by moving files. A coding benchmark score jumps, but new private tasks do not asks what a jump on a public coding benchmark says about a model. Here the interview problem is the data lineage and split construction that made a held-out support answer appear new when it was not.