Model and Inference Engineering · Principal
The converted checkpoint loads. Why does RoPE change the model's answers?
The question
Interview question
A team converts a trained checkpoint to a new inference runtime. Every Q and K tensor has the expected shape, the model produces fluent text, and short smoke tests pass. Longer prompts are noticeably worse. The converter copied the projection weights and enabled the runtime's rotary embedding kernel. What was missing?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
RoPE rotates pairs of coordinates in each attention head by an angle determined by position. Implementations can lay out those paired coordinates differently. One may pair adjacent coordinates. Another may pair corresponding coordinates from two halves of the head. A conversion that changes the layout of activations must permute the corresponding Q and K projection rows, or use a rotary kernel matching the original layout. Copying bytes into an identically shaped tensor is not enough. Hugging Face's weight conversion documentation explicitly has a PermuteForRope operation for this kind of conversion.
I would choose a fixed token sequence and compare source and destination implementations layer by layer, before looking at generated prose. First compare the unrotated Q and K in a common coordinate convention. Then compare rotated Q and K at positions zero, one and a much later position. Finally compare per-head attention logits, outputs and the next-token logits. At position zero some errors can be less visible, so a one-token smoke test can give false comfort. Use deliberately different values in every coordinate and head. A fixture with repeated values can hide a permutation.
The exact conversion depends on the architecture, rotary dimension, head layout, tensor-parallel slicing and any fused QKV representation. V does not get the same rotary permutation merely because it sits next to Q and K in a fused tensor. Likewise, changing the frequency base or position IDs is a separate possible fault. If Q/K parity holds before rotation and diverges after it, inspect the kernel and position handling. If it already diverges before rotation, inspect the projection weight conversion. This gives a fault boundary instead of guessing from a benchmark score.
What if the converted model passes a perplexity check? I would ask which lengths and whether the source and destination were evaluated on identical token IDs and masks. Put logit parity tests at several positions and across a cached decode step into the release gate. A small numerical tolerance is reasonable across kernels, but a systematic head-coordinate permutation is not numerical noise. The GQA cache has the right shape. Why do answers change after tensor parallel reshaping? is about mapping query heads to KV heads after reshaping. Here the head mapping can be correct while the coordinate pairing inside each head is wrong.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →