Model and Inference Engineering · Principal
The pipeline updated a stage before backward reached it. Which weights made the gradient?
The question
Interview question
A custom pipeline trainer overlaps microbatches and optimizer updates. Microbatch A runs its forward pass through stage 2 using weights `W0`. Before A's backward reaches that stage, another microbatch finishes and stage 2 updates to `W1`. Backward for A now uses the live parameter storage. Is A's gradient still the derivative of its forward computation?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Generally, no. A gradient for an operation such as y = Wx depends on the computation that actually ran. Backpropagating through A as though it used W1 can produce the wrong input gradient for the previous stage, even if A's activation and labels were saved. In many ordinary autograd paths, an in-place version check may catch a modification. A custom fused or distributed path can evade that protection. The question is about the schedule's parameter-version contract, not only whether an exception is raised.
One safe synchronous pattern runs the pipeline's microbatches through forward and backward under the same weights, accumulates their gradients, then steps the optimizer at a flush boundary. That costs bubble time or memory depending on the schedule. An asynchronous schedule can update earlier, but it needs to retain the appropriate version for each microbatch's backward. PipeDream's pipeline-parallel training work describes parameter versioning for this reason. Stashing versions addresses the forward/backward mismatch at a stage. It does not by itself promise that an asynchronous pipeline has identical optimization semantics to a fully synchronous batch.
I would trace one microbatch with a version tag at every stage's forward and backward, and log when each optimizer update becomes visible. Freeze a tiny two-stage model and compare gradients with a sequential reference under the schedule we claim to implement. Include activation recomputation, since a recomputed forward under W1 while the original forward used W0 silently changes the graph unless it restores the correct version. Also check tied parameters shared across stages, which make independent stage updates even harder to reason about.
If the interviewer says 1F1B schedules always have this bug, I would push back. Many practical 1F1B implementations accumulate a synchronous global step and update only after a flush. The issue appears when an optimizer update becomes visible inside an outstanding forward/backward lifetime. Ask for the version trace, not just the schedule's name.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →