Model and Inference Engineering · Staff
Only two MoE experts run per token. Why does serving still need room for all the experts?
The question
Interview question
A team compares an MoE model with a dense model using their active parameter counts. Two experts run for each token, so they estimate that its weight memory will look like the active count too. The replica cannot fit. Where did the estimate go wrong?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Active parameters describe the work selected for one token, not the set of weights a serving fleet must make available. The router can send the next token to another expert. Different tokens in one batch can visit many experts, and each layer has its own experts. If you want predictable latency for all permitted inputs, those weights must be resident somewhere sufficiently close to compute, or you need a deliberate loading policy with a measured miss cost. Switch Transformers makes the basic sparsity distinction clear. It grows total parameter count while keeping computation per token more controlled. It does not say the whole model suddenly needs only the active weights in storage.
I would make a memory budget from the checkpoint, not the marketing number. Count bytes for all resident expert weights, shared attention and embeddings, scales or metadata for the chosen quantization, runtime workspace, communication buffers and KV cache under the intended concurrency. Then show the mapping across GPUs. Replicating every expert on each GPU is very different from sharding experts across GPUs. The latter saves per-device weight memory, but routed tokens now travel to their assigned device. Interconnect bandwidth, all-to-all behavior, load imbalance and slow experts enter the latency budget. The aggregate fleet still has to hold the weights, perhaps with replication for throughput and availability.
Could we keep cold experts in CPU RAM or storage? Yes, for a workload where routing is skewed enough and misses fit the latency target. I would measure expert selection by layer and tenant, bursts, co-occurrence within batches, warmup after deploy, and the p95 and p99 for a cold route. A cache sized for yesterday's popular experts can thrash after the distribution changes. You cannot infer the correct cache from the average two-experts-per-token figure. Also distinguish expert weight memory from expert activation and KV memory. KV scales with live sequence length and concurrency even though the MoE FFN routes sparsely.
Someone may argue that a token never touches most experts, so all of them need not be on the same GPU. Correct. Expert parallelism is exactly the design option to examine. But the router still needs a reachable home for every expert it might select, and the token may pay to get there. Compare a dense and MoE system at equal task quality and the actual batch, not equal active FLOPs alone. I would ship the topology only after measuring memory headroom, communication, queueing under skew and recovery when one expert-hosting GPU fails. That is the real serving cost of sparse activation.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →