Model and Inference Engineering · Staff
The MoE has plenty of total capacity. Why are tokens still dropped?
The question
Interview question
A sparse mixture of experts layer has eight experts. The capacity budget across all eight exceeds the batch's token count, yet training traces show overflow and a quality drop for a small language. Someone suggests adding up all expert slots and concludes the router must be broken. How do you explain the drop, and what would you change without hiding the problem in an average?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Imagine 800 routed token assignments and eight experts with 120 slots each. Total slots are 960. That does not help if 300 assignments choose expert 2 and its local capacity is 120. The other experts cannot automatically lend it their empty slots. With top two routing, count assignments, not just source tokens, because one token may consume capacity at two experts. In many top-k implementations the exact overflow path is configuration dependent. Some assignments may be dropped, sent through a residual path, or rerouted. I would inspect this implementation before saying the token was lost entirely.
Sparse routing lets the model have more parameters without applying every expert to every token. The router's preference can become concentrated because certain token types or domains look similar to it. Expert capacity is usually local to an expert and often a device or expert parallel group. Switch Transformer describes capacity and load balancing in a top-one design. Its specific policy is an example, not the behavior of every MoE. A dashboard with average expert utilization can conceal one full expert beside seven underused ones, and even an aggregate drop rate can conceal almost all drops landing on one language.
I would get the routing trace by layer, expert, batch, sequence position and evaluation cohort. Record assigned tokens, accepted tokens, overflow policy, router probabilities and auxiliary balancing loss. Check whether the small language is overrepresented in a few hot experts or whether it gets displaced by other tokens competing in the same batches. Compare per-domain task loss and quality with and without overflow. Also inspect which experts are hot at the same time across the expert parallel group. A rank-level average of GPU utilization does not tell us the per expert dispatch queue or communication cost.
There are a few knobs, and none is free. Raise capacity factor and you spend memory and compute on larger expert buffers. Improve load balancing and you may trade off useful specialization against a more even distribution. Change the router, add an overflow expert, alter batch construction, or change the number of experts and parallel placement. A second choice expert helps only if that route is allowed and has room. Do not claim “no drops” by silently ignoring the overloaded assignment in the metric. Measure tokens that took the actual path and the resulting quality.
Suppose the interviewer says the small language is only one percent of traffic, so the overall quality is stable. That is exactly why I would keep a named slice. If a routed layer's overflow is clustered, the average can look excellent while a legitimate cohort fails repeatedly. I would test an intervention against that cohort, common tasks, throughput and tail latency. It may turn out the router is usefully specializing and only the capacity policy is wrong. It may turn out the training mixture makes one expert a universal shortcut. Both require more than summing eight slot counts.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →