Look at what the cache stores, not just its shape. With rotary position embeddings, the query at position i is rotated by i and the key at position j by j. Their dot product then carries the relative phase between those positions. RoFormer gives the underlying construction. One valid cache convention stores keys after rotation. Another stores raw projected keys and rotates them when read. Both can produce the right attention scores. Mixing them cannot. If prefill stored a rotated key and decode rotates it again on retrieval, the key has effectively been given position 2j while the query still has position i.

Why might prefill tests pass? A full prefill path can compute attention with freshly rotated Q and K and never exercise the cache-reader conversion that decode uses. Even the first generated token may look plausible if a common prefix masks the error. Compare a full forward pass over the prompt plus one extra token with a cached prefill followed by one-token decode, layer by layer. Use more than a single position. Position zero hides a second rotation because its angle is zero. Check short and long positions, odd cache block boundaries, and a nonzero prefix offset. The first divergent attention logits are more useful than a final text comparison.

I would make the cache contract explicit in the type or metadata: raw versus rotated keys, the position indices and RoPE scaling configuration used, and whether any fused attention kernel applies rotation internally. On append, a new token must follow the same convention as the prefetched tokens. On transfer between prefill and decode workers, both ends need the same contract. A sliding window can keep logical positions increasing even when physical slots wrap, so slot number must not quietly become the rotation index. Then remove one rotation at the correct boundary and prove parity on a stored checkpoint across representative context lengths. A broad change that “rotates all cached keys for consistency” could break the prefill-to-decode handoff again.