Understanding Key-Value (KV) Cache Scaling
Why auto-regressive text generation transforms transformer inference from compute-bound to memory-capacity-bound.
⚡ What is the KV Cache in Transformer Models?
During auto-regressive generation, each newly generated token attends to all prior tokens in the sequence. Without caching, the model would recompute key ($K$) and value ($V$) vectors for every prior token at every generation step—resulting in $O(N^2)$ redundant compute overhead.
By caching the Key and Value matrices in GPU High-Bandwidth Memory (HBM), the model computes keys and values only for the current new token, reducing generation to $O(N)$ operations. However, this optimization shifts the bottleneck directly onto GPU VRAM capacity.
🔄 Evolution: MHA vs GQA vs MLA
1. Multi-Head Attention (MHA) — e.g. LLaMA 1, GPT-3
Every query head possesses its own dedicated key and value head ($H_{kv} = H_q$). At long context lengths (32k+), the KV cache rapidly balloons to hundreds of gigabytes, making high-concurrency production serving financially prohibitive.
2. Grouped-Query Attention (GQA) — e.g. Llama 3, Mistral, Qwen 2.5
Query heads are grouped into clusters that share a single key-value head (such as 8 query heads per 1 KV head in Llama 3 70B). This slashes KV cache memory by up to 8x (87.5% reduction) with negligible degradation in reasoning capabilities.
3. Multi-Head Latent Attention (MLA) — e.g. DeepSeek-V2 / V3
DeepSeek projects keys and values into a compressed low-rank latent space ($d_c = 512$) along with decoupled RoPE positional embeddings ($d_R = 64$). This compresses the KV cache by up to 93% compared to traditional MHA, allowing massive 128k context windows on modest GPU clusters.
🗜️ KV Cache Quantization: FP8 & INT4
Modern inference engines like vLLM, SGLang, and TensorRT-LLM support quantizing the cached KV tensors:
- FP16 / BF16 (16-bit): 2 bytes per element. Baseline precision with zero accuracy degradation.
- FP8 (8-bit E4M3): 1 byte per element. Cuts KV memory by exactly 50% with <0.1 perplexity impact on modern GPUs with native FP8 tensor cores (Ada Lovelace, Hopper).
- INT4 (4-bit): 0.5 bytes per element. Cuts KV memory by 75%, allowing massive concurrency on single RTX 4090 or A100 GPUs.
🔒 Zero-Data Collection Guarantee
KVCacheCalc executes 100% locally in your client browser. Your architecture sizing parameters, concurrency estimates, and custom model dimensions are never transmitted to external servers or logged in remote datastores.