Model and Inference Engineering · Staff
RoPE scaling accepts 128K tokens. Why does the model lose track near the end?
The question
Interview question
A model trained mostly on short sequences is served with a rotary position scaling configuration and a 128K context limit. It passes a single passkey test at 120K, but when asked to compare a late exception with an early rule in a long policy document, it confidently applies the old rule. Is this a position encoding failure? What would you measure before promising a 128K product feature?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would first separate three claims that the context number tends to collapse. The server can allocate and process 128K tokens. The attention mechanism assigns positional information at those indices. The model can use the right evidence across that span for our tasks. A successful API call proves the first. A passkey needle proves something narrower than the third.
In RoPE, query and key components are rotated as a function of token position. Their attention score therefore carries relative position information. Extending the indices beyond the range seen during training is not just increasing a buffer size. Position interpolation, for example, maps a larger position range back into the old one, then fine tunes the model for the longer sequences. The Position Interpolation paper reports useful long context results for that particular recipe. It does not establish that a model with an arbitrary scaling factor understands every dependency at 128K.
The failure might not be positional at all. Was the late exception actually in the tokenized input after the system prompt, attachments and truncation? Did the document parser preserve the word “unless” and its section boundary? Is the answer driven by a more frequent early rule, or by a retrieval step that never surfaced the exception? Check the exact token offsets, the sequence of transformed text and the model's answer with the late paragraph moved near the question. If moving it fixes the answer, we have evidence of a distance or salience problem, but still not a proof of a particular RoPE defect. Move the early rule as well. Otherwise we may be measuring a preference for whichever statement occurs last.
For a clean probe I would create pairs that require combining two facts rather than copying one marker. Hold wording and total length constant, then vary the positions of the rule, exception and question independently. Include contradictory distractors, multiple plausible exceptions and sections that look similar. Score whether the model names the applicable exception and cites the correct source span, not just whether it repeats a keyword. Plot accuracy by both evidence positions and separation. Test the short original window at the same time, because a scaling or fine tuning change can regress short context behavior. Use a held out set of real documents alongside synthetic probes.
There is a systems cost too. A model can technically accept 128K and still have unacceptable prefill latency, KV cache pressure, batching interference or price. I would report quality and serving cost together for the proposed distribution of prompt lengths. If the interviewer asks whether to fix this by retrieving only a few chunks, that may be the right product design for some tasks. It can also throw away the exception we needed. Evaluate retrieval recall for the decisive span and compare a narrower, structured evidence input against full context on the same cases.
I would advertise a 128K input limit only as that. For a claim that it answers long document questions reliably, I want position-sensitive task results, a known training and scaling recipe, and a failure policy when the relevant evidence cannot be established. The model accepting the tokens is the start of that argument.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →