KV Cache Optimization

A Study Guide
Why the Cache Grows  •  MQA / GQA / MLA  •  Quantization (KIVI · KVQuant · QJL · TurboQuant)  •  Token Eviction  •  Cross-Layer Sharing  •  PagedAttention

Contents

  1. Why the KV Cache Is the Bottleneck
  2. Fewer KV Heads — MQA, GQA, MLA
  3. Quantization — Fewer Bits per Entry
  4. Token Eviction & Sparsity
  5. Cross-Layer Sharing
  6. Systems & Memory Management
  7. Spotlight: TurboQuant
  8. Combining Methods
  9. References
Part 1

Why the KV Cache Is the Bottleneck

What is the KV cache?

A transformer generates text one token at a time. At each step, every layer's attention needs the keys (K) and values (V) of all previous tokens. Rather than recompute them every step, they are computed once and stored — the KV cache. This makes generation fast, but the cache must be held in memory for the entire sequence.

Inference only?
Yes — the KV cache exists only during autoregressive generation (inference). Training processes the whole sequence in parallel (one forward pass, causal mask), so there are no sequential steps to cache across. Two caveats: the architectural methods below (GQA/MQA/MLA, cross-layer) are baked into the model and so are present during training; and training loops that generate (e.g. RLHF rollouts) use the cache during their generation phase.
Figure 1.1 — Prefill builds the cache; decode reuses and extends it
Prefill (process the prompt once) Prompt: 4 tokens processed in parallel→ compute K,V for each, store in cache K,V t1 K,V t2 K,V t3 K,V t4 Decode (one new token at a time) New token attends to ALL cached K,Vthen appends its own K,V to the cache t1 t2 t3 t4 t5 ✚ The cache grows by one K,V entry per generated token — forever, until the sequence ends. Every decode step must read the entire cache from memory → bandwidth, not compute, is the limiter. Two costs: (1) MEMORY — the cache must fit in RAM/VRAM alongside the weights. (2) BANDWIDTH — it is re-read every single step; for long context this dominates latency.

The size formula

KV bytes = 2 (K&V) × layers × kv_heads × head_dim × seq_len × batch × bytes_per_value

The cache scales linearly with sequence length and batch size. A worked example for a 7B-class model (32 layers, 32 KV heads, head_dim 128, FP16):

2 × 32 × 32 × 128 × 4096 × 1 × 2 bytes ≈ 2.1 GB (at 4k context, batch 1)

At 32k context that is ~17 GB — often larger than the model weights themselves. This is why every term in the formula is a lever, and the optimization families below each attack a different one.

FamilyTerm it shrinksLossy?Examples
Fewer KV headskv_headsDesigned-in (trained)MQA, GQA, MLA
Quantizationbytes_per_valueMildly lossyKIVI, KVQuant, QJL, TurboQuant
Token evictionseq_len (kept)LossyStreamingLLM, H2O, SnapKV
Cross-layer sharinglayersDesigned-inCLA, YOCO
Memory managementwaste / fragmentationLosslessPagedAttention, offloading
Part 2

Fewer KV Heads — MQA, GQA, MLA

The idea

Standard multi-head attention (MHA) stores a separate K and V for every attention head. But many query heads can share the same keys and values with little quality loss. Sharing fewer K/V heads shrinks the cache directly — and because it is baked into the architecture, it is free at inference time.

Figure 2.1 — MHA → GQA → MQA: how many query heads share each K/V
MHA1 K/V per query head 4 query heads 4 K/V heads cache ×4 GQAgroups share K/V 4 query heads 2 K/V heads cache ×2 MQAall share one K/V 4 query heads 1 K/V head cache ×1 GQA is the modern default — near-MHA quality at a fraction of the cache. Used in Llama-2/3, Mistral, Qwen. MQA is the most aggressive (one K/V), but can lose a little quality and stability. MLA (Multi-head Latent Attention, DeepSeek-V2/V3) goes further It compresses K/V into a small shared low-rank latent vector, cached instead of full K/V — a large reduction.
MethodK/V headsCache vs MHANotes
MHA= query heads1× (baseline)Best quality, biggest cache
GQAa few groups~2–8× smallerThe standard trade-off
MQA1up to ~#heads smallerMost aggressive; small quality risk
MLAlow-rank latentlarge reductionCache a compressed latent, not raw K/V
Part 3

Quantization — Fewer Bits per Entry

The idea

Instead of storing each cached number in 16-bit float, store it in 8, 4, 3, or even 2 bits. This is the same principle as weight quantization, but applied to the cache — and it has a twist: keys and values behave differently, so the best methods treat them separately.

Figure 3.1 — The key insight: per-channel for keys, per-token for values
KEYS — quantize per channel Some channels have large outliers → scale each channel (column) on its own each column its own scale VALUES — quantize per token Values are smoother across channels → scale each token (row) on its own each row its own scale

This asymmetric treatment (the core of KIVI) lets keys go to 2 bits without the outlier channels wrecking accuracy. Other methods add tricks: quantizing keys before rotary position encoding (KVQuant), or using a random projection so no calibration data is needed (QJL).

MethodBitsKey trickCalibration?
KIVI2-bitPer-channel keys, per-token values (asymmetric)Tuning-free
KVQuant2–4 bitPre-RoPE + per-channel keys + non-uniform + outlier sparsityLight calibration
QJL~1–4 bitQuantized Johnson–Lindenstrauss random projection on keysNone (data-free)
TurboQuant~3-bit K / 2-bit VPolar rotation + 1-bit residual correction (see Part 7)None (random rotations)
Why "asymmetric"?
Keys have a few extreme outlier channels; values are smoother. Quantizing each along its sensitive axis is the single biggest reason 2-bit KV works at all.
Part 4

Token Eviction & Sparsity

The idea

Not every past token matters for the next one. Eviction methods keep only the K/V of "important" tokens and drop the rest, so the cache stops growing without bound. The trick is deciding which tokens to keep.

Figure 4.1 — What each eviction policy keeps (green) vs drops (gray)
StreamingLLM — attention sinks + recent window first fewkeep the first tokens ("sinks") + a sliding recent window; drop the middle. recent H2O — keep the "heavy hitters" (highest attention) score tokens by accumulated attention; keep the top scorers wherever they are. SnapKV — compress the prompt at prefill using attention patterns a recent "observation window" votes for which prompt tokens to keep — one-shot at prefill. Trade-off: eviction is lossy — a dropped token is gone. Safe when attention is naturally sparse (most long-context tasks).
MethodPolicyWhen it decides
StreamingLLMKeep first ("sink") tokens + recent windowStreaming, fixed rule
H2OKeep high cumulative-attention "heavy hitters"During decode
SnapKVCluster + select important prompt tokensOnce, at prefill
PyramidKVSmaller budgets in deeper layers (info funnels)Per-layer budget
Part 5

Cross-Layer Sharing

The idea

The size formula has a layers term. Adjacent layers often produce similar keys and values — so why store a separate cache for each? Cross-layer sharing caches K/V once and reuses it across several layers, cutting the cache by the sharing factor.

  • CLA (Cross-Layer Attention): groups of layers share one K/V cache. A 2× sharing factor roughly halves the cache with minimal quality change.
  • YOCO (You Only Cache Once): a decoder-decoder design where the entire model caches K/V one time in a global module and every later layer reads it — a large reduction, especially for very long context.
Relation to MLA
MLA shrinks the per-layer cache (low-rank latent); cross-layer sharing shrinks the number of caches. They are orthogonal and can stack.
Part 6

Systems & Memory Management

The idea

These methods don't make the cache smaller — they stop you from wasting the memory you have. Classic serving pre-allocates a big contiguous block per request for the maximum possible length; most of it sits empty. That waste limits how many requests fit on a GPU.

Figure 6.1 — PagedAttention: KV cache in small pages, like OS virtual memory
Naïve: one big contiguous block per request used reserved but empty (wasted) internal fragmentation → fewer requests fit PagedAttention: many small pages, allocated on demand a page table maps logical positions → physical pages; near-zero waste Result (vLLM): much higher batch size / throughput — and pages can be SHARED across requests e.g. a shared system prompt is stored once and pointed to by many sequences (prefix sharing). Offloading (FlexGen) When the cache still won't fit, move cold K/V to CPU RAM or disk and stream it back as needed.
Part 7 — Spotlight

TurboQuant

TurboQuant (Google Research, ICLR 2026) is a recent KV-cache quantization method that pushes to ~3-bit keys / 2-bit values — about 6× compression — with near-zero accuracy loss and, notably, no calibration. It is a strong example of the quantization family from Part 3.

Figure 7.1 — TurboQuant pipeline: rotate, quantize, correct
K / V vectorFP16 ① Polar rotationrandom rotation →evens out the values ② Quantizeto ~3-bit (keys)/ 2-bit (values) ③ 1-bit residual fixa tiny QJL correction termrecovers the rounding error Random (not learned) rotations → calibration-free, drop-in on any model, at any time. ~6× memory reduction · up to ~8× faster attention on H100 · within ~2.7× of the theoretical (Shannon) optimum.

Why it stands out

  • Calibration-free. Because the rotations are random rather than fitted to data, there is no calibration pass — you can apply it to any model immediately. (QJL shares this property; many older methods need calibration data.)
  • Provably near-optimal. The paper derives a Shannon lower bound on the best possible distortion for any quantizer at a given bit budget, and shows TurboQuant sits within a small constant factor (~2.7×) of it — so there is limited headroom left for a better method at the same bits.
  • Practical. It ships with GPU (Triton) kernels and a vLLM integration path, so the compression turns into real speedups, not just smaller memory.
Where it fits
TurboQuant is one method in the quantization family (Part 3) — it does not replace GQA, eviction, or PagedAttention. In practice you would stack it on top of them.
Part 8 — Capstone

Combining Methods

They stack — each attacks a different term

Just like general model compression, KV-cache methods are orthogonal and multiply when combined. A long-context serving stack typically layers all four families:

Figure 8.1 — A production long-context stack
ArchitectureGQA / MLA QuantizeTurboQuant / KIVI EvictSnapKV / H2O ServePagedAttention (vLLM) Fewer heads × fewer bits × fewer tokens × zero waste = order-of-magnitude smaller cache, longer context, higher batch.
Your situationReach for
Designing/choosing a modelGQA (or MLA) — biggest win, free at inference
Memory-bound, can't retrainKV quantization — KIVI / QJL / TurboQuant (drop-in)
Very long prompts, sparse attentionEviction — SnapKV / H2O / StreamingLLM
Serving many concurrent requestsPagedAttention (vLLM) + prefix sharing
Cache still won't fitOffloading (FlexGen) — to CPU/disk
On-device / edge note
For edge inference (and vision-language models, whose image tokens make sequences long before any text), the cache often rivals the weights. Calibration-free quantization (TurboQuant/QJL) plus GQA is the highest-leverage, lowest-friction combination.
References

Foundational Papers & Further Reading

Fewer KV Heads

  1. Shazeer, N. (2019). Fast Transformer Decoding: One Write-Head is All You Need (MQA). arxiv.org/abs/1911.02150
  2. Ainslie, J. et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. EMNLP 2023. arxiv.org/abs/2305.13245
  3. DeepSeek-AI (2024). DeepSeek-V2 (introduces Multi-head Latent Attention, MLA). arxiv.org/abs/2405.04434

Quantization

  1. Liu, Z. et al. (2024). KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache. ICML 2024. arxiv.org/abs/2402.02750
  2. Hooper, C. et al. (2024). KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization. NeurIPS 2024. arxiv.org/abs/2401.18079
  3. Zandieh, A. et al. (2024). QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead. arxiv.org/abs/2406.03482
  4. Google Research (2026). TurboQuant: Near-Optimal Vector Quantization for KV Cache Compression. ICLR 2026. arxiv.org/abs/2504.19874

Token Eviction & Sparsity

  1. Xiao, G. et al. (2023). Efficient Streaming Language Models with Attention Sinks (StreamingLLM). ICLR 2024. arxiv.org/abs/2309.17453
  2. Zhang, Z. et al. (2023). H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models. NeurIPS 2023. arxiv.org/abs/2306.14048
  3. Li, Y. et al. (2024). SnapKV: LLM Knows What You are Looking for Before Generation. NeurIPS 2024. arxiv.org/abs/2404.14469
  4. Cai, Z. et al. (2024). PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling. arxiv.org/abs/2406.02069

Cross-Layer Sharing

  1. Brandon, W. et al. (2024). Reducing Transformer Key-Value Cache Size with Cross-Layer Attention (CLA). arxiv.org/abs/2405.12981
  2. Sun, Y. et al. (2024). You Only Cache Once: Decoder-Decoder Architectures for Language Models (YOCO). NeurIPS 2024. arxiv.org/abs/2405.05254

Systems & Memory Management

  1. Kwon, W. et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM). SOSP 2023. arxiv.org/abs/2309.06180
  2. Sheng, Y. et al. (2023). FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. ICML 2023. arxiv.org/abs/2303.06865