A transformer generates text one token at a time. At each step, every layer's attention needs the keys (K) and values (V) of all previous tokens. Rather than recompute them every step, they are computed once and stored — the KV cache. This makes generation fast, but the cache must be held in memory for the entire sequence.
The cache scales linearly with sequence length and batch size. A worked example for a 7B-class model (32 layers, 32 KV heads, head_dim 128, FP16):
At 32k context that is ~17 GB — often larger than the model weights themselves. This is why every term in the formula is a lever, and the optimization families below each attack a different one.
| Family | Term it shrinks | Lossy? | Examples |
|---|---|---|---|
| Fewer KV heads | kv_heads | Designed-in (trained) | MQA, GQA, MLA |
| Quantization | bytes_per_value | Mildly lossy | KIVI, KVQuant, QJL, TurboQuant |
| Token eviction | seq_len (kept) | Lossy | StreamingLLM, H2O, SnapKV |
| Cross-layer sharing | layers | Designed-in | CLA, YOCO |
| Memory management | waste / fragmentation | Lossless | PagedAttention, offloading |
Standard multi-head attention (MHA) stores a separate K and V for every attention head. But many query heads can share the same keys and values with little quality loss. Sharing fewer K/V heads shrinks the cache directly — and because it is baked into the architecture, it is free at inference time.
| Method | K/V heads | Cache vs MHA | Notes |
|---|---|---|---|
| MHA | = query heads | 1× (baseline) | Best quality, biggest cache |
| GQA | a few groups | ~2–8× smaller | The standard trade-off |
| MQA | 1 | up to ~#heads smaller | Most aggressive; small quality risk |
| MLA | low-rank latent | large reduction | Cache a compressed latent, not raw K/V |
Instead of storing each cached number in 16-bit float, store it in 8, 4, 3, or even 2 bits. This is the same principle as weight quantization, but applied to the cache — and it has a twist: keys and values behave differently, so the best methods treat them separately.
This asymmetric treatment (the core of KIVI) lets keys go to 2 bits without the outlier channels wrecking accuracy. Other methods add tricks: quantizing keys before rotary position encoding (KVQuant), or using a random projection so no calibration data is needed (QJL).
| Method | Bits | Key trick | Calibration? |
|---|---|---|---|
| KIVI | 2-bit | Per-channel keys, per-token values (asymmetric) | Tuning-free |
| KVQuant | 2–4 bit | Pre-RoPE + per-channel keys + non-uniform + outlier sparsity | Light calibration |
| QJL | ~1–4 bit | Quantized Johnson–Lindenstrauss random projection on keys | None (data-free) |
| TurboQuant | ~3-bit K / 2-bit V | Polar rotation + 1-bit residual correction (see Part 7) | None (random rotations) |
Not every past token matters for the next one. Eviction methods keep only the K/V of "important" tokens and drop the rest, so the cache stops growing without bound. The trick is deciding which tokens to keep.
| Method | Policy | When it decides |
|---|---|---|
| StreamingLLM | Keep first ("sink") tokens + recent window | Streaming, fixed rule |
| H2O | Keep high cumulative-attention "heavy hitters" | During decode |
| SnapKV | Cluster + select important prompt tokens | Once, at prefill |
| PyramidKV | Smaller budgets in deeper layers (info funnels) | Per-layer budget |
The size formula has a layers term. Adjacent layers often produce similar keys and values — so why store a separate cache for each? Cross-layer sharing caches K/V once and reuses it across several layers, cutting the cache by the sharing factor.
These methods don't make the cache smaller — they stop you from wasting the memory you have. Classic serving pre-allocates a big contiguous block per request for the maximum possible length; most of it sits empty. That waste limits how many requests fit on a GPU.
TurboQuant (Google Research, ICLR 2026) is a recent KV-cache quantization method that pushes to ~3-bit keys / 2-bit values — about 6× compression — with near-zero accuracy loss and, notably, no calibration. It is a strong example of the quantization family from Part 3.
Just like general model compression, KV-cache methods are orthogonal and multiply when combined. A long-context serving stack typically layers all four families:
| Your situation | Reach for |
|---|---|
| Designing/choosing a model | GQA (or MLA) — biggest win, free at inference |
| Memory-bound, can't retrain | KV quantization — KIVI / QJL / TurboQuant (drop-in) |
| Very long prompts, sparse attention | Eviction — SnapKV / H2O / StreamingLLM |
| Serving many concurrent requests | PagedAttention (vLLM) + prefix sharing |
| Cache still won't fit | Offloading (FlexGen) — to CPU/disk |