Security, Governance and Platform · Principal
The prefix cache never returns another tenant's text. What can latency still reveal?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Two tenants share an inference fleet. Tenant A repeatedly sends a confidential standard prompt. Tenant B can submit candidate prefixes and measure first-token latency. When B guesses a cached prefix, prefill work is skipped and the response tends to be faster. The cache did not return A's text. But the timing difference may tell B that a particular prefix was recently used.
This is a membership side channel, not the same as serving A's answer to B. It is most useful to an attacker when the secret is guessable from a small set, such as a known document template or a yes/no candidate. Network jitter can make one measurement noisy, but repeated measurements and controlled traffic may expose a difference. vLLM's security guidance explicitly discusses prefix-cache timing attacks in shared deployments and recommends a secret cache_salt at the tenant boundary. This answer describes a fleet that omitted an effective isolation value, not a universal vulnerability in every prefix cache.
I would trace the actual cache key and trust boundary. Does identity include exact token IDs and the needed model configuration? Is a tenant-scoped unpredictable salt mixed in before any reusable prefix block can collide across tenants? An account ID passed in clear is an identifier, not necessarily a secret salt. Bind the salt at the trusted gateway, not from a user-supplied field that B can copy. Make cache sharing policy explicit: sharing within a tenant may be worth the cost savings, while cross-tenant sharing of private prefixes should be disabled unless the data is truly public and the risk accepted.
How do we establish the leak? Run controlled A/B timing probes with and without A's prefix cached, across realistic network conditions. Measure distribution of time to first token and the attacker's classification accuracy, not a single impressive hit. Also check whether the cache's metrics, error messages or tokens-used accounting reveal a hit more directly. Inspect all prefix and multimodal cache layers, because isolating one hash table does not isolate a second cache keyed by user-provided media IDs.
The pushback is cost. Per-tenant salts reduce cross-tenant hits and may raise prefill work. Measure that tradeoff against the confidentiality boundary. A shared public system prompt can be handled under a separate public cache policy if the architecture supports it, without merging private user prefixes. A semantic cache served tenant A's private answer to tenant B. What now? concerns a semantic answer cache that directly returns another tenant's private answer. Here no foreign answer is returned. The observable is whether a candidate prompt prefix was present in a shared inference cache.
Continue reading
Related questions
Read beyond the question
Explore more security, governance and platform
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →