Data and Knowledge Systems · Principal
A deleted document still appears in the next training run. Which copy did the loader read?
The question
Interview question
A source document is removed from the active corpus on Monday. A new model run starts Tuesday from what the team calls a fresh dataset, yet a sampled batch contains its text. The loader reads pretokenized shards from object storage and some workers keep a local cache. How do you find the copy and make a future-run exclusion reliable?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
“Removed from the active corpus” identifies one boundary. It says little about a shard built last week, a copied dataset snapshot, a local Arrow cache, or a worker that already opened a file. A training run reads the material named by its manifest and data loader, not a live view of the source UI. Hugging Face Datasets documents separate download and processed Arrow caches. Other training stacks make different choices, but materialized derivatives are a normal part of a high-throughput pipeline. I would trace the actual bytes the Tuesday run read.
Start with a stable source document ID and revision, then find every derived chunk, normalized text record, tokenized sequence and packed shard that contains it. A document can span shards, and a packed sequence can mix several sources. Keep a lineage map from source IDs to derived object IDs and byte or token ranges. For the suspect batch, record run ID, manifest digest, shard digest, worker, sample index and token offset. If an exact text match is ambiguous, compare normalized spans and upstream identifiers rather than claiming a source based on a short phrase. The Tuesday run may have pinned a Monday snapshot that legitimately still named the document under the old recipe. The product name “fresh” is not proof that a new manifest was made.
For future runs, create a new immutable dataset generation after applying the exclusion to every derived representation. Regenerate affected shards or use a tombstone filter whose semantics are verified at the loader boundary. Publish a manifest that contains only approved shard digests, pin each training job to that manifest, and make the run record it. Validate by scanning the resolved token stream or a representative complete index of source lineage, not just the source table. Then restart workers from the new generation. A cache is safe only if its key includes the relevant source and transformation version and the job cannot silently reuse an old generation. Megatron Core's dataset notes illustrate how cached sample and shuffle indices can be separate from source datasets. That is one concrete place to check, not a universal cache bug.
What about a job already running? Define the cutover requirement. If exclusion must take effect before the next optimizer step, pause the job, establish which batches have already been consumed, rebuild the allowed stream and resume from a checkpoint with a clearly new data generation. Merely deleting the object while workers hold it open is unreliable and can make recovery worse. If the requirement is only for newly launched runs, pinning a validated new manifest may be enough. An already trained checkpoint may have incorporated the content. Removing the document from future datasets does not reverse past weight updates.
I would close the incident with two proofs: no forbidden source IDs in the new loader's realized stream, and an inventory of runs and checkpoints that did ingest them. What happens to those prior models is a separate product and governance decision based on the actual exclusion requirement. The engineering part begins by naming every copy and the exact data generation each run used.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →