Model and Inference Engineering · Staff
The checkpoint loads after a SwiGLU kernel rewrite. Why did model quality fall?
The question
Interview question
An inference team fuses a model's gated MLP for speed. Shapes match, weight checksums match, and the model emits fluent text, but quality drops. The new kernel splits a packed projection into two halves and treats the former `up` half as the gate. Could that matter when both halves have the same shape and are multiplied together?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Yes. In a SwiGLU-style block, the operation is roughly down(SiLU(gate(x)) * up(x)). The activation is applied to one branch before multiplication. Switching the trained gate and up weights computes SiLU(up(x)) * gate(x). Multiplication is commutative, but applying SiLU to only one operand is not. The weight files can load cleanly because the matrices have matching dimensions. NVIDIA's Transformer Engine Llama example shows the gate, up and down projections and where the activation belongs. The exact packing order is a checkpoint and kernel contract, not something to guess from a tensor shape.
I would compare the reference and fused path at one layer on a fixed hidden-state tensor. Check gate projection, up projection, activated gate, product and down projection before comparing whole-model logits. That localizes the first divergence. Inspect how the converter concatenated the weights, their transpose convention, any per-channel quantization scales and the fused kernel's split index. A correct formula with scales attached to the wrong half can make the same symptom. Use several inputs, including negative gate values where SiLU and identity differ substantially. A test on all-positive values can accidentally hide a swap.
Then run end-to-end deterministic prompts on the same checkpoint and tokenizer, with the same numerical precision and batching shape. Expect some tolerance for reduced-precision kernels, but an abrupt, systematic logit shift after one layer is not ordinary rounding. Compare task slices that stress exact choice, not just whether text looks grammatical. Use checksum or manifest metadata to assert projection naming and ordering at conversion, and a golden layer output test for every supported architecture variant. A fused kernel for another model family may pack branches differently.
The interviewer may say the loss remains finite. Of course. The wrong MLP still maps vectors to vectors, and a deep model may produce fluent answers despite degraded internal features. The attention scores use model width instead of head width. What changes? concerns a temperature change in attention. This question is about a non-equivalent algebraic rewrite that survives shape and load checks. The appropriate rollback condition is evidence of semantic divergence against the reference, not a runtime crash.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →