The identity of a fetch is not the identity of its text. Byte-identical copies can have different URLs. Near duplicates can differ in a footer, timestamp, markup, language boilerplate or a few words. Repeated substrings can survive even if no two whole documents match. Research on deduplicating language-model training data finds substantial near-duplicate material and repetitive spans, and studies effects on memorization and evaluation overlap. It does not mean every repeated phrase in this hypothetical run was memorized because of duplication. We would need to measure that.

I would audit at several granularities. Normalize obvious crawl artifacts, then compare exact content hashes, approximate document similarity and repeated long substrings. Keep provenance so we know which URLs or source collections contributed each copy. Do the analysis before mixing and sampling, because repeated documents may also be oversampled by the data recipe. Measure multiplicity per source and domain, not only a global duplicate percentage. A small number of widely syndicated articles can contribute a large fraction of tokens on a narrow topic.

Then decide the dedup policy with the training objective in mind. Remove exact copies aggressively, use near-duplicate clustering with manual review near thresholds, and preserve legitimate revisions and distinct documents that happen to share boilerplate. Do not blindly collapse every two code files with similar scaffolding. A dedup pass can erase minority languages, niche documentation or repeated but important examples if it uses a crude similarity measure. Evaluate rare-domain coverage and downstream quality alongside memorization and contamination risk.

If the interviewer asks whether this is just a benchmark-leak issue, no. Train/eval overlap is one consequence, but repeated text also changes the effective token mixture and can encourage regurgitation. Split held-out data by content similarity, not URL alone, then track the source-cluster IDs through train and evaluation. Test the pipeline with a mirrored page, a small revision, a copied paragraph inside a new page, and a genuinely independent article using a common template. Train and eval have different document IDs. Why did the model see the test answers? covers test answers appearing in training despite different document IDs. This page asks whether deduplication actually controls repetition within pretraining, even before an evaluation set enters the picture.