It depends on what “step” means in the schedule, but a recipe defined by successful optimizer updates has drifted. The learning rate may warm up or decay without the corresponding weight update. An early run of overflows can exhaust a meaningful part of warmup before training has taken those steps. PyTorch's AMP examples explain that GradScaler may skip optimizer.step() if it finds nonfinite gradients, and the optimizer scheduler guidance addresses optimizer and scheduler call ordering. Calling a scaler wrapper is not proof that an optimizer update occurred.

I would log at least three counters: batches consumed, backward or accumulation windows attempted, and committed optimizer updates. Add skipped-overflow counts, effective tokens used for each update, current scale and learning rate. Trace which counter feeds the scheduler, checkpoint naming, evaluation cadence and resume logic. In distributed training, all ranks must agree whether a step committed, and any schedule state must advance consistently. One training rank has nonfinite gradients while the global loss looks normal. Can the step commit? asks whether a nonfinite distributed step should commit. Here we are examining all the other state that may advance even when it did not.

If the intended schedule is tokens seen rather than successful updates, advancing after a skipped optimizer step may be an explicit choice. Then say so and check whether tokens that produced no update should count toward that budget. Schedules based on wall time, tokens and optimizer updates answer different questions. A training team should not move between them by accident when changing AMP behavior. For a committed-update schedule, call the scheduler only when the optimizer actually ran. Use an authoritative success signal supported by the framework, rather than guessing from a scaler value that may change for other reasons across versions.

How do we test it? Force a controlled nonfinite gradient in one accumulation window, verify that model and optimizer state do not change, and verify the scheduler and successful-step counter behave according to the chosen recipe. Resume from a checkpoint after that skipped window and compare the subsequent learning rates and updates. A loss curve alone may not reveal this, since many runs tolerate small schedule differences. The bug matters most during warmup, frequent overflows, or tight token budgets where skipped steps are not rare.