Active parameters describe the work selected for one token, not the set of weights a serving fleet must make available. The router can send the next token to another expert. Different tokens in one batch can visit many experts, and each layer has its own experts. If you want predictable latency for all permitted inputs, those weights must be resident somewhere sufficiently close to compute, or you need a deliberate loading policy with a measured miss cost. Switch Transformers makes the basic sparsity distinction clear. It grows total parameter count while keeping computation per token more controlled. It does not say the whole model suddenly needs only the active weights in storage.

I would make a memory budget from the checkpoint, not the marketing number. Count bytes for all resident expert weights, shared attention and embeddings, scales or metadata for the chosen quantization, runtime workspace, communication buffers and KV cache under the intended concurrency. Then show the mapping across GPUs. Replicating every expert on each GPU is very different from sharding experts across GPUs. The latter saves per-device weight memory, but routed tokens now travel to their assigned device. Interconnect bandwidth, all-to-all behavior, load imbalance and slow experts enter the latency budget. The aggregate fleet still has to hold the weights, perhaps with replication for throughput and availability.

Could we keep cold experts in CPU RAM or storage? Yes, for a workload where routing is skewed enough and misses fit the latency target. I would measure expert selection by layer and tenant, bursts, co-occurrence within batches, warmup after deploy, and the p95 and p99 for a cold route. A cache sized for yesterday's popular experts can thrash after the distribution changes. You cannot infer the correct cache from the average two-experts-per-token figure. Also distinguish expert weight memory from expert activation and KV memory. KV scales with live sequence length and concurrency even though the MoE FFN routes sparsely.

Someone may argue that a token never touches most experts, so all of them need not be on the same GPU. Correct. Expert parallelism is exactly the design option to examine. But the router still needs a reachable home for every expert it might select, and the token may pay to get there. Compare a dense and MoE system at equal task quality and the actual batch, not equal active FLOPs alone. I would ship the topology only after measuring memory headroom, communication, queueing under skew and recovery when one expert-hosting GPU fails. That is the real serving cost of sparse activation.