Model and Inference Engineering · Staff
The learning rate decays at the planned step. Why has the model seen half the intended tokens?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An optimizer step is not a fixed amount of data unless the number of valid training tokens per step is fixed. The recipe says warm up over 1,000 steps, assuming eight million nonpadding tokens in that interval. A new packing and filtering pipeline delivers only four million valid targets by step 1,000. The learning-rate scheduler is behaving exactly as configured. The schedule in token exposure has changed. Hugging Face's optimizer schedules take a number of warmup and total steps, while its trainer state tracks steps and can track input tokens separately.
I would count three quantities, since they answer different questions: raw tokens read, input positions processed, and nonignored target tokens that contributed to loss. Padding and masked prompt spans can make them diverge sharply. Compare these per optimizer update and cumulatively across both runs, along with sequence lengths, gradient accumulation, world size and skipped optimizer steps. If the job uses a custom token-based scheduler, verify what it calls a token and when it advances its counter.
The fix depends on the intended recipe. If the scientific comparison is by processed training tokens, define warmup, decay and checkpoints in that unit, and make the conversion to steps explicit for the actual batch distribution. If the optimizer update count itself is the controlled variable, keep the step schedule but stop claiming equivalent token exposure. Changing tokens per update may also change gradient noise and task mixture, so matching one counter alone does not make the runs identical.
Could we simply double the number of warmup steps after noticing the gap? For a fresh run, that is a testable recipe change. Halfway through a run, it changes the remaining learning-rate path and needs a new experiment record, not a silent patch. The optimizer skipped a step. Why did the learning rate schedule advance? covers a scheduler advancing when an optimizer step was skipped. This question assumes steps really happened, but each step contained fewer useful tokens than the schedule designer assumed.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →