Model and Inference Engineering · Staff
The MoE picks the same two experts. Why did changing router normalization change the answer?
The question
Interview question
An inference rewrite produces the same top two expert IDs for every test token. The team calls it equivalent. Quality drops anyway. Where would you look?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
At the weights used to combine those experts. Imagine four router logits whose full softmax assigns probabilities 0.45, 0.35, 0.15, 0.05. The selected experts are the first two. If you keep their full-softmax values, their contributions sum to 0.80. If you softmax only the selected logits, their weights become 0.5625 and 0.4375, summing to one. The expert IDs match, but the mixture output is scaled and potentially altered at every MoE layer. A residual path can carry that difference through the network. A learned output scale or other routing convention changes the exact arithmetic again.
This is why I would inspect the checkpoint's actual router contract, not assume there is one universal MoE formula. Does it apply softmax before or after top-k? Are selected weights renormalized? Is there a scaling factor? Is it sigmoid routing, grouped selection, shared experts, or a capacity rule that changes which tokens reach an expert? Megatron Core exposes pre-softmax routing and top-k scaling as separate configuration choices, and its router API distinguishes selection scores from expert-combination probabilities. Those knobs are part of the model, not a serving preference to guess at.
Compare the old and new paths on frozen hidden states. Record router logits, chosen indices, combination weights, per-expert outputs and the summed layer output. Include ties and near ties, where small precision changes can flip selection, and measure a complete forward pass before blaming sampling. If the weights differ while the expert IDs agree, you already have a counterexample to equivalence. If both match, move further downstream.
The interviewer might ask whether renormalizing is always wrong. No. Some models are trained that way. The mistake is serving a checkpoint under a different rule than the one it learned. The MoE has plenty of total capacity. Why are tokens still dropped? asks why capacity drops tokens during training. This one has no dropped tokens at all. It tests whether two implementations compute the same model.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →