RoPE rotates pairs of coordinates in each attention head by an angle determined by position. Implementations can lay out those paired coordinates differently. One may pair adjacent coordinates. Another may pair corresponding coordinates from two halves of the head. A conversion that changes the layout of activations must permute the corresponding Q and K projection rows, or use a rotary kernel matching the original layout. Copying bytes into an identically shaped tensor is not enough. Hugging Face's weight conversion documentation explicitly has a PermuteForRope operation for this kind of conversion.

I would choose a fixed token sequence and compare source and destination implementations layer by layer, before looking at generated prose. First compare the unrotated Q and K in a common coordinate convention. Then compare rotated Q and K at positions zero, one and a much later position. Finally compare per-head attention logits, outputs and the next-token logits. At position zero some errors can be less visible, so a one-token smoke test can give false comfort. Use deliberately different values in every coordinate and head. A fixture with repeated values can hide a permutation.

The exact conversion depends on the architecture, rotary dimension, head layout, tensor-parallel slicing and any fused QKV representation. V does not get the same rotary permutation merely because it sits next to Q and K in a fused tensor. Likewise, changing the frequency base or position IDs is a separate possible fault. If Q/K parity holds before rotation and diverges after it, inspect the kernel and position handling. If it already diverges before rotation, inspect the projection weight conversion. This gives a fault boundary instead of guessing from a benchmark score.

What if the converted model passes a perplexity check? I would ask which lengths and whether the source and destination were evaluated on identical token IDs and masks. Put logit parity tests at several positions and across a cached decode step into the release gate. A small numerical tolerance is reasonable across kernels, but a systematic head-coordinate permutation is not numerical noise. The GQA cache has the right shape. Why do answers change after tensor parallel reshaping? is about mapping query heads to KV heads after reshaping. Here the head mapping can be correct while the coordinate pairing inside each head is wrong.