🧠
KVCacheCalc LLM VRAM Sizing Engine
AI Systems Engineering • Hardware Architecture

Understanding Key-Value (KV) Cache Scaling

Why auto-regressive text generation transforms transformer inference from compute-bound to memory-capacity-bound.

⚡ What is the KV Cache in Transformer Models?

During auto-regressive generation, each newly generated token attends to all prior tokens in the sequence. Without caching, the model would recompute key ($K$) and value ($V$) vectors for every prior token at every generation step—resulting in $O(N^2)$ redundant compute overhead.

By caching the Key and Value matrices in GPU High-Bandwidth Memory (HBM), the model computes keys and values only for the current new token, reducing generation to $O(N)$ operations. However, this optimization shifts the bottleneck directly onto GPU VRAM capacity.

🔄 Evolution: MHA vs GQA vs MLA

1. Multi-Head Attention (MHA) — e.g. LLaMA 1, GPT-3

Every query head possesses its own dedicated key and value head ($H_{kv} = H_q$). At long context lengths (32k+), the KV cache rapidly balloons to hundreds of gigabytes, making high-concurrency production serving financially prohibitive.

2. Grouped-Query Attention (GQA) — e.g. Llama 3, Mistral, Qwen 2.5

Query heads are grouped into clusters that share a single key-value head (such as 8 query heads per 1 KV head in Llama 3 70B). This slashes KV cache memory by up to 8x (87.5% reduction) with negligible degradation in reasoning capabilities.

3. Multi-Head Latent Attention (MLA) — e.g. DeepSeek-V2 / V3

DeepSeek projects keys and values into a compressed low-rank latent space ($d_c = 512$) along with decoupled RoPE positional embeddings ($d_R = 64$). This compresses the KV cache by up to 93% compared to traditional MHA, allowing massive 128k context windows on modest GPU clusters.

🗜️ KV Cache Quantization: FP8 & INT4

Modern inference engines like vLLM, SGLang, and TensorRT-LLM support quantizing the cached KV tensors:

  • FP16 / BF16 (16-bit): 2 bytes per element. Baseline precision with zero accuracy degradation.
  • FP8 (8-bit E4M3): 1 byte per element. Cuts KV memory by exactly 50% with <0.1 perplexity impact on modern GPUs with native FP8 tensor cores (Ada Lovelace, Hopper).
  • INT4 (4-bit): 0.5 bytes per element. Cuts KV memory by 75%, allowing massive concurrency on single RTX 4090 or A100 GPUs.

🔒 Zero-Data Collection Guarantee

KVCacheCalc executes 100% locally in your client browser. Your architecture sizing parameters, concurrency estimates, and custom model dimensions are never transmitted to external servers or logged in remote datastores.