In multi-head latent attention, the content part of each token's key and value can be represented by a smaller shared latent vector. Instead of keeping full per-head K and V for every past token, the server keeps that compressed content state. The query and output projections can be arranged to work with the latent without reconstructing every full key and value during decode. DeepSeek-V2 introduces this design, and vLLM's MLA backend notes describe the cached latent and separate positional key in its implementation.

Position is the complication. Rotary position encoding applies a position-dependent transformation to query and key components. A cached content latent alone does not carry the required position-specific key term in the form this architecture uses. DeepSeek's decoupled RoPE keeps a smaller positional key component alongside the content latent. The query compares content through its compressed path and position through this separate rotated path. So "one latent vector" is a useful shorthand, but the actual cached entry includes positional state as well.

I would make a candidate derive what is stored per token in this checkpoint and serving backend: latent width, positional-key width, number of layers, dtype and any extra metadata. Then compare with ordinary MHA or GQA KV bytes. A claim like "MLA reduces the entire KV cache by exactly X percent" cannot be copied across models with different dimensions. It also does not mean attention compute, weight storage or prefill cost fall by the same fraction.

If we remove the positional key to save a little more, we are changing the trained attention function. A test with repeated content at two positions is a good probe: can the query still distinguish where the matching token appeared? Check output parity between the reference implementation and compressed-cache decode, especially after a long prefix. Grouped query attention cuts KV memory. Why might prefill still be expensive? asks why GQA's smaller KV does not erase prefill cost. This question shows how MLA reduces per-token cached content while preserving a separate path for positional information.