LLM KV Cache
Memory Size Calculator
Calculate exact GPU VRAM needed for Key-Value Caches, model weights, and concurrent batch sizes. Size your deployment across RTX 4090, A100, and H100 clusters.
1. Model & Hardware Parameters
Required GPU VRAM
GPU Hardware Compatibility Matrix
Which GPUs can host this configuration with zero Out-Of-Memory (OOM) errors.
The Mathematics of KV Cache Scaling
How Key-Value caching trades GPU high-bandwidth memory (HBM) for extreme inference throughput.
📐 Multi-Head Attention (MHA / GQA) Formula
- 2: Separate Key (K) and Value (V) tensors
- L: Transformer layers (e.g., 32 for 8B, 80 for 70B)
- H_kv: Number of Key-Value heads (8 in GQA)
- d_h: Head dimension ($d_{\text{model}} / H_{\text{query}}$, typically 128)
- P: Precision in bytes (2 for FP16, 1 for FP8, 0.5 for INT4)
- C: Context length in tokens; B: Batch size
⚡ DeepSeek Multi-Head Latent Attention (MLA)
DeepSeek-V2 and V3 compress the Key-Value cache into a low-rank latent vector of dimension $d_c = 512$ plus decoupled RoPE dimension $d_R = 64$. This compresses KV cache memory by up to 93% compared to standard MHA, allowing 128k context on massive batch sizes.
Frequently Asked Questions
Essential engineering questions on KV cache allocation.
What is the formula for calculating LLM KV Cache size? ▼
For standard Multi-Head Attention (MHA) and Grouped-Query Attention (GQA), the formula is: KV Cache (bytes) = 2 × n_layers × n_kv_heads × d_head × precision_bytes × context_length × batch_size. The factor of 2 accounts for storing both Key and Value tensors. In FP16 precision, precision_bytes = 2; in FP8 it is 1, and in INT4 it is 0.5.
Why does the KV cache exceed model weights at large context windows? ▼
Model weights are static and do not scale with request length. However, the KV cache grows linearly with both context length and batch size. For example, running Llama-3-70B with FP16 KV cache across 32 concurrent streams at 64k context requires over 130 GB of VRAM just for the KV cache alone, vastly eclipsing the model weights.
How does Grouped-Query Attention (GQA) reduce KV cache memory? ▼
Standard Multi-Head Attention pairs every query head with its own key and value head. Grouped-Query Attention (GQA), used in models like Llama 3, Mistral, and Qwen, shares a single key-value head among multiple query heads (e.g., an 8:1 ratio in Llama 3 70B), reducing KV cache memory consumption by up to 8x with minimal impact on accuracy.
What is FP8 and INT4 KV cache quantization? ▼
Inference engines like vLLM, TensorRT-LLM, and SGLang support quantizing KV cache tensors to FP8 (8-bit floating point, cutting memory by 50%) or INT4 (4-bit integer, cutting memory by 75%). This allows higher batch sizes and longer context windows to fit onto smaller GPUs like the RTX 4090 or single A100.