Sparsity reduces expert arithmetic, not the cost of finding and reaching those experts. With expert parallelism, the router assigns tokens to experts resident on other devices. Their hidden states must be dispatched, processed, and sent back for the outputs to be combined in the original token order. Research on distributed MoE serving discusses how expert-parallel all-to-all communication can bound inference efficiency. A small decode batch has fewer tokens with which to amortize dispatch, synchronization and network startup. One hot destination or slow rank can hold the layer even when average compute is low.

I would profile an MoE layer, not just the model: router time, token counts by expert and destination, dispatch bytes, all-to-all latency, expert compute, return transfer, combine time and the slowest rank. Compare prefill and decode separately. If the slowdown is only in decode, grouping more requests may improve communication efficiency but hurt queueing or inter-token SLOs. Replicating hot experts or placing frequently co-selected experts near tokens might reduce traffic, but the router distribution changes with workload and may change with model revision. Test with real traffic, not a synthetic uniform expert assignment.

The placement decision has several constraints. Keep enough expert replicas for capacity and failure recovery. Avoid pushing all hot routes onto one device. If experts are sharded, account for weight communication as well as token dispatch. If replicas are used, decide how their weights are loaded, updated and versioned. Quantizing experts can reduce memory and change quality. Collocating everything may remove network costs and exceed HBM. The optimum is measured for the topology and traffic mix, not inferred from “top two.”

An interviewer may ask whether the router is simply imbalanced. It might be, and One MoE expert group is hot while GPU averages look fine already covers a hot expert group hidden by GPU averages. This question remains even with reasonably balanced destinations: an all-to-all round trip and synchronization can cost more than the saved expert compute at low batch size. Only two MoE experts run per token. Why does serving still need room for all the experts? asks why every expert must still be available somewhere. Here we decide where they live and whether moving tokens to them was worthwhile.