Routing is only half the operation. An expert-parallel MoE layer groups token representations by destination expert, sends them to the right devices, runs expert networks, and then restores outputs to the original token positions. For top-two routing, it may also combine two outputs with the right weights. A correct top-k expert list does not prove that the return trip preserved token identity. Megatron Core's token-dispatcher documentation explicitly describes permutation, all-to-all communication, and unpermutation.

Imagine input positions A, B and C routed in that order, but the transport groups B and C at one device and A at another. If the combine path uses the pre-permutation index to scatter returned rows, A can receive B's transformed representation. Shapes and collective sizes all match. Average expert utilization looks good. The model's final answer is wrong because the residual stream at one position now contains another token's information. If top-two outputs are summed, also check that probabilities and expert output rows are associated with the same original token.

I would give each input token an artificial, easily recognized vector and use experts with simple known transformations. Compare a single-device reference to expert-parallel execution for uneven routes, multiple tokens to one expert, empty experts, top-two routing, capacity overflow and different batch sizes. Log a compact token index and chosen expert through dispatch and combine in a debug build. Check the inverse permutation and counts across ranks, including the case where an optimization fuses sorting with communication.

The pushback is that training loss would surely explode. Sometimes it would, but a small mistake affecting only one route shape can hide in average loss and surface under a rare production batch. Gate a dispatcher change on layer-output parity and task slices, not only throughput. The MoE has plenty of total capacity. Why are tokens still dropped? asks why a hot expert drops assignments. This question assumes the intended experts ran and tests whether their outputs reached the intended positions.