The serving fleet is rolling from weights v12 to v13. A conversation had already run prefill for its history on v12. The next turn lands on v13, but the router reuses a remote KV handle because the token prefix matches. The request succeeds and produces an odd answer. There may be no exception because tensor shapes can match across revisions.

KV is not a cached copy of text. Each attention layer's keys and values were computed from that layer's activations under particular weights, tokenizer behavior, positional scheme, adapter and cache layout. A new model cannot in general continue decoding from KV produced by different weights, even if every token ID is identical. vLLM's prefix-caching design describes reuse of previously computed KV for shared prefixes and the cache identity inputs it handles. The cross-revision remote-cache error in this question is a hypothetical deployment bug, not a claim about vLLM's default behavior.

My first move is to log the model revision on the request, prefill worker, KV artifact and decode worker, then prove where they diverge. A model alias such as production is not a revision. Use an immutable digest of the effective weights and relevant serving configuration in the cache namespace or artifact descriptor. Check it at the consumer, even when the router already checked it. Tokenizer revision, adapter identity, attention implementation compatibility, cache dtype and layout may also be material. Different engines can establish a narrower compatibility contract, but the default should be fail closed on a mismatch.

How do we avoid dropping every active conversation during a rollout? The durable conversation is the token or message history, not the KV handle. Pin an in-flight generation to one compatible revision. For the next turn, either route to the old revision while it is available or replay the full accepted history through v13 to rebuild KV. That costs prefill work and may change the answer because it is a new model, which is an expected product choice that should be explicit. Draining old workers is a capacity and latency decision. It is never permission to reinterpret old KV under new weights.

Test a cache hit across a rolling deploy with revisions that have identical shapes, a LoRA change, and a failed old worker. Compare the new revision's cached continuation with a fresh v13 prefill, then audit cross-revision hit counters. Prefill and decode disagree about the KV cache concerns mismatched prefill and decode workers within a request. This case is a stale persisted handle crossing a model upgrade between turns.