Model and Inference Engineering · Principal
All MoE experts are equally busy. Why did answer quality get worse?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Equal traffic is an operational goal, not proof that the router chose the best experts for each token. The router wants to send a token to experts that help predict it. The serving system wants to avoid hot experts, dropped tokens and stragglers. A strong balance penalty can move tokens away from a specialist merely to fill a quiet expert. Utilization looks beautiful while the main prediction objective suffers. DeepSeek-V3's technical report describes auxiliary-loss-free balancing partly to avoid interference from an auxiliary balancing loss. That motivates the investigation, not a claim that every balance loss harms quality.
First separate where the regression entered. Check training loss excluding auxiliary terms, task evals by domain, expert routing by token type and load by expert at both training and serving time. If the reported training loss includes the balance penalty, a lower total number can conceal a worse language-model loss. If training quality is unchanged but serving quality drops, inspect capacity, dispatch and dropped-token behavior under real batches. The MoE has plenty of total capacity. Why are tokens still dropped? handles capacity-driven drops. Here assume no tokens were dropped and ask whether equalizing the route itself changed which computation each token received.
A useful experiment holds model, data and compute budget fixed while varying the balance pressure. Measure prediction loss and task quality alongside worst-expert load, tail latency and capacity overflows. Look for groups that need specialized behavior, such as code tokens, rare languages or math notation. If a rare group loses its preferred route, aggregate accuracy may barely move while those tasks fail. Balance also has a time scale. Uniform counts across a day can hide a single hot request batch that still stalls the GPU.
Would I remove balancing entirely? Only if the unbalanced model fits the available serving system and the resulting tail cost is acceptable. A different balancing method, expert placement or capacity plan may preserve quality without an unbounded hot spot. One MoE expert group is hot while GPU averages look fine covers a hot expert group in serving. This question is the training and routing tradeoff underneath that operational symptom: did we optimize the distribution of work at the expense of the work itself?
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →