Model and Inference Engineering · Principal
The crawl has unique URLs. Why does pretraining keep seeing the same article?
The question
Interview question
A pretraining corpus is deduplicated by URL and document ID. Training logs still show a technical article appearing hundreds of times through mirrors, tracking parameters, copied forum posts and chunks cut from larger pages. The model begins reproducing long phrases from it. Why did the dedup check pass?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The identity of a fetch is not the identity of its text. Byte-identical copies can have different URLs. Near duplicates can differ in a footer, timestamp, markup, language boilerplate or a few words. Repeated substrings can survive even if no two whole documents match. Research on deduplicating language-model training data finds substantial near-duplicate material and repetitive spans, and studies effects on memorization and evaluation overlap. It does not mean every repeated phrase in this hypothetical run was memorized because of duplication. We would need to measure that.
I would audit at several granularities. Normalize obvious crawl artifacts, then compare exact content hashes, approximate document similarity and repeated long substrings. Keep provenance so we know which URLs or source collections contributed each copy. Do the analysis before mixing and sampling, because repeated documents may also be oversampled by the data recipe. Measure multiplicity per source and domain, not only a global duplicate percentage. A small number of widely syndicated articles can contribute a large fraction of tokens on a narrow topic.
Then decide the dedup policy with the training objective in mind. Remove exact copies aggressively, use near-duplicate clustering with manual review near thresholds, and preserve legitimate revisions and distinct documents that happen to share boilerplate. Do not blindly collapse every two code files with similar scaffolding. A dedup pass can erase minority languages, niche documentation or repeated but important examples if it uses a crude similarity measure. Evaluate rare-domain coverage and downstream quality alongside memorization and contamination risk.
If the interviewer asks whether this is just a benchmark-leak issue, no. Train/eval overlap is one consequence, but repeated text also changes the effective token mixture and can encourage regurgitation. Split held-out data by content similarity, not URL alone, then track the source-cluster IDs through train and evaluation. Test the pipeline with a mirrored page, a small revision, a copied paragraph inside a new page, and a genuinely independent article using a common template. Train and eval have different document IDs. Why did the model see the test answers? covers test answers appearing in training despite different document IDs. This page asks whether deduplication actually controls repetition within pretraining, even before an evaluation set enters the picture.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →