Yes, quite easily. Twenty percent of documents selected is not twenty percent of tokens optimized. Suppose a code document averages 4,000 retained tokens and a non-code document averages 1,000. At an 80 to 20 document split, code contributes 800,000 tokens for every 800,000 non-code tokens over 1,000 selected documents. That is 50% of retained tokens before truncation, repetition, packing or masking. The exact numbers are made up for this question, but the denominator mismatch is real.

I would ask for the full chain from source to gradient. What unit does the mixer draw, a document, a fixed-length sequence, an already packed sample, or a dataset shard? At what stage do we label a token as code? How much does filtering remove from each source? Are long documents split into many samples, short ones packed with neighbors, or examples repeated after a small domain is exhausted? Are padding, prompt tokens or masked labels counted in the dashboard? In a standard fixed-length GPT training dataset, a 20% draw of equal length samples can be close to 20% of input tokens. It is not legitimate to use the document-length example to explain that setup without evidence. Megatron Core's dataset description makes the fixed sequence length and blending index choices explicit. Read the actual sampler contract.

The review number I want is effective contribution to the objective, along with the units used to calculate it. For each source and cohort, count selected documents, retained tokens, sampled occurrences, input tokens, loss-bearing tokens and total loss weight after any per-token or per-example weighting. Deduplicate cautiously for an additional view of unique material. A 47% code share could come from legitimate long code documents, repeated code shards, a bug in mixture weights, or simply a token classifier that calls every Markdown file with a code fence “code.” The raw source percentage alone cannot distinguish them.

I would freeze a manifest of source revisions, filtering code, tokenizer, sampling seed, mixture weights and data loader configuration, then inspect actual batches from many ranks and training steps. The intended distribution at configuration time is one thing. The realized distribution after worker partitioning, cache reuse and epoch rollover is another. Reconcile the counters to the optimizer's global step and total weighted target tokens, with a clear rule for mixed documents. This is also where an accidental repeated shard shows up.

Should we force it back to 20% token share? Only if that was the intended objective and the quality evidence supports it. Maybe code was deliberately oversampled because it is valuable. DoReMi is one example of research that treats domain weights as an optimization choice rather than the raw size of the corpora. I would compare model quality on code, general language and important rare domains under controlled mixtures. The point is to know which quantity we chose. A slide saying “20% code” cannot stand in for a measured training stream.