Model and Inference Engineering · Staff
We replaced GELU with SwiGLU at the same FFN width. Why did the model get larger?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Because SwiGLU is not a drop-in activation inside the same two-matrix feed-forward block. An ordinary block projects from model width d to hidden width h, applies an activation, then projects back to d. Ignoring bias, that is 2dh weights. SwiGLU makes two separate projections from d to h, gates one with a smooth activation and multiplies their outputs, then projects back. That is 3dh weights. The original GLU variants paper describes this three-matrix construction and reduces hidden width when comparing parameter budgets.
If the old FFN used h = 4d, it had about 8d² weights. Reusing h = 4d for SwiGLU produces 12d², a 50 percent increase in the FFN weights, not a 50 percent increase in the whole Transformer. To match the old FFN count, choose h = 8d/3 before any hardware-friendly rounding. Biases, activation storage, fused kernels and exact model conventions change the detailed cost, but they do not erase the extra projection. This is why an architecture spec needs the intermediate dimension, not just the name of its activation.
The serving impact is not only file size. During low-batch decode, more weight bytes may have to be read, and a different projection layout can change kernel efficiency. During training, parameter gradients and optimizer states scale too. I would compare checkpoint parameter count, layer-wise bytes, decode latency at relevant batches, prefill throughput and quality at matched training compute or parameter budget. If quality improves after adding 50 percent more FFN weights, we have not isolated the benefit of gating.
Could we take a trained GELU checkpoint and simply insert a new gate? No. The new projection has no learned weights with the intended function. It needs training or an explicitly validated conversion. The checkpoint loads after a SwiGLU kernel rewrite. Why did model quality fall? is about a SwiGLU kernel rewrite changing a model that already uses SwiGLU. This question asks how to size the architecture fairly before training and serving it.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →