Model and Inference Engineering · Principal
The LoRA adapter changed. Why did a cached prefix still act like the old one?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
KV entries are the result of running a particular prefix through a particular effective model. A LoRA adapter can change the projections that produce those keys and values. The same token IDs under a new adapter version do not imply the same cached KV.
Imagine a custom deployment that publishes revised adapter bytes under a stable logical ID, such as tenant-42. A prefix cache key includes token IDs and that logical LoRA ID, but not the adapter version. After the serving fleet swaps the adapter, an old cache block can still match. The request then combines old-version prefix activations with new-version computation on the uncached suffix. It may return a fluent but inconsistent answer, and ordinary single-request tests that start with a cold cache will miss it. vLLM's prefix-caching design explicitly includes LoRA IDs among extra hash inputs. That establishes the need to distinguish adapters, while a same-ID hot replacement is the hypothetical integration failure here. Do not infer that vLLM necessarily makes this mistake in its own adapter lifecycle.
The cache identity needs to include an immutable identity for everything that changes the effective forward pass, or the deployment must invalidate dependent blocks before serving the new version. A versioned adapter ID is one option. An epoch attached to the loaded adapter and prefix-cache namespace is another. Check base-model version too. Merely changing a routing label without changing the key is not enough.
To prove the fault, compare cold and warm requests with identical tokens before and after an adapter update. Record the adapter artifact digest, loaded version on each replica, prefix block key and cache-hit decision. Disable caching for one canary while keeping the new adapter, then compare its logits or targeted behavior with the warm path. A difference only on warm hits points to cache identity or invalidation. Mixed-version replicas can be a separate cause, so pin both requests to the same replica for the first experiment.
If someone asks why not purge the entire cache on every update, that is a safe initial rollout choice for a small fleet. At scale it can hurt first-token latency for unrelated tenants. Version-scoped keys or targeted invalidation preserve reuse without letting activations from one effective model masquerade as another.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →