Two tenants share an inference fleet. Tenant A repeatedly sends a confidential standard prompt. Tenant B can submit candidate prefixes and measure first-token latency. When B guesses a cached prefix, prefill work is skipped and the response tends to be faster. The cache did not return A's text. But the timing difference may tell B that a particular prefix was recently used.

This is a membership side channel, not the same as serving A's answer to B. It is most useful to an attacker when the secret is guessable from a small set, such as a known document template or a yes/no candidate. Network jitter can make one measurement noisy, but repeated measurements and controlled traffic may expose a difference. vLLM's security guidance explicitly discusses prefix-cache timing attacks in shared deployments and recommends a secret cache_salt at the tenant boundary. This answer describes a fleet that omitted an effective isolation value, not a universal vulnerability in every prefix cache.

I would trace the actual cache key and trust boundary. Does identity include exact token IDs and the needed model configuration? Is a tenant-scoped unpredictable salt mixed in before any reusable prefix block can collide across tenants? An account ID passed in clear is an identifier, not necessarily a secret salt. Bind the salt at the trusted gateway, not from a user-supplied field that B can copy. Make cache sharing policy explicit: sharing within a tenant may be worth the cost savings, while cross-tenant sharing of private prefixes should be disabled unless the data is truly public and the risk accepted.

How do we establish the leak? Run controlled A/B timing probes with and without A's prefix cached, across realistic network conditions. Measure distribution of time to first token and the attacker's classification accuracy, not a single impressive hit. Also check whether the cache's metrics, error messages or tokens-used accounting reveal a hit more directly. Inspect all prefix and multimodal cache layers, because isolating one hash table does not isolate a second cache keyed by user-provided media IDs.

The pushback is cost. Per-tenant salts reduce cross-tenant hits and may raise prefill work. Measure that tradeoff against the confidentiality boundary. A shared public system prompt can be handled under a separate public cache policy if the architecture supports it, without merging private user prefixes. A semantic cache served tenant A's private answer to tenant B. What now? concerns a semantic answer cache that directly returns another tenant's private answer. Here no foreign answer is returned. The observable is whether a candidate prompt prefix was present in a shared inference cache.