Suppose an 8,000-token prompt runs correctly in one prefill. Split it into four 2,000-token chunks, reuse the KV cache between chunks, and the answer changes. Chunking should change when work is scheduled, not which token position the model thinks it is processing. I would inspect the position IDs and cache offsets at each boundary.

For a RoPE-based model, keys and queries are rotated according to position. The first token of the second chunk is token 2,000 in the full prompt, not token zero. If custom prefill code resets its local position counter for each chunk, it rotates new keys and queries as if a new sequence started. Those keys then sit beside cached keys produced under the original coordinate system. Attention is operating on a history that no single unchunked forward pass would produce. Transformers' cache guidance explains that custom generation must keep cache positions accurate when appending to cached states.

There is another boundary to check. A token in the new chunk must see the permitted old prefix plus earlier tokens in its own chunk, but never future tokens in that chunk. A mask built only for a local 2,000 by 2,000 block can omit the old prefix or accidentally expose the future. Distinguish that from a RoPE error by comparing the position IDs, mask and attention logits at the first token after each split.

I would run one short deterministic prompt in one piece and with several chunk sizes, including a split immediately before the token the model must retrieve. Compare per-layer hidden states or logits at corresponding absolute token positions. Check that the cache length grows by exactly the number of real tokens, not pad slots, and that the final decode token starts at the right position. A chunk-size-dependent change with all else fixed is strong evidence of a serving implementation error, though floating-point kernel differences can also make near-tie outputs diverge. Compare numeric differences before calling every changed token a bug.

An interviewer may ask whether RoPE is relative, so a reset should cancel out. It cancels a common shift applied to both query and key. Here old cached keys keep their earlier rotation while new keys and queries restart. That is not a common shift of the whole sequence.