Model and Inference Engineering · Staff
The sliding-window KV cache fits. Why did evicting every layer break long prompts?
The question
Interview question
An inference server sees a model with sliding-window attention and caps each layer's KV cache at the window size. Memory drops sharply. Then answers needing information from the start of a long prompt regress. An engineer says older tokens were outside the window anyway. Is that true for this checkpoint?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Check the layer configuration, not the model's shorthand label. Some architectures combine local-window layers with global-attention layers, attention sinks, or another mechanism for retaining longer-range information. A layer that only attends to its recent window can use a rolling KV store under its defined mask. A global layer cannot discard keys and values merely because a neighboring layer uses a window. Hugging Face's cache guide says cache growth stops at the window for layers using sliding-window attention. That qualifier matters. Mistral 7B describes a rolling buffer for its sliding-window design, but its configuration is not a license to rewrite every other model's cache policy.
I would inspect per-layer attention types and the serving engine's mapping from layer to cache class. Compare reference logits for a short prompt, at the first wraparound, and several windows later. A bug that appears exactly when a ring buffer overwrites slot zero points to eviction, position mapping or an attention mask that still addresses the old slot. Record logical token positions separately from physical ring slots. Check prefill chunks and resumed decoding, not only one uninterrupted generation. If the model mixes local and global layers, profile the memory separately for each type and confirm that global layers retain what their mask can reach.
There is also a boundary within a genuinely local layer. A token outside the current window cannot be attended to directly by that layer, but earlier information may have flowed through retained hidden states in other positions. That is a property of the trained architecture, not a reason to change the cache mask. Conversely, keeping every old KV in local layers wastes memory but may not change output if the mask excludes them. The server needs to implement the checkpoint's exact attention and position semantics.
For release, I would use long-range retrieval probes at several distances, run logit comparisons across the eviction point, and measure memory per live token at realistic concurrency. The claim “the KV cache fits” is only meaningful after checking correctness. Swap KV to host memory or recompute it? weighs swapping KV versus recomputing. This question asks which KV is safe to throw away in the first place.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →