🧠
KVCacheCalc LLM VRAM Sizing Engine
Supports MHA, GQA & DeepSeek MLA • vLLM & SGLang Compatible

LLM KV Cache
Memory Size Calculator

Calculate exact GPU VRAM needed for Key-Value Caches, model weights, and concurrent batch sizes. Size your deployment across RTX 4090, A100, and H100 clusters.

Llama 3.1 8B KV Rate 128 KB Per token (FP16 GQA)
Llama 3.3 70B KV Rate 320 KB Per token (FP16 GQA)
FP8 KV Cache Savings 50% Cut vLLM & TensorRT-LLM
DeepSeek MLA Mode ~93% Cut Multi-Head Latent Attn
Advertisement
Responsive AI Hardware & Cloud GPU Infrastructure Ad Slot

1. Model & Hardware Parameters

16,384 tokens
2k 8k 32k 128k 1M
8 streams
1 (Local/Solo) 16 (Prod Server) 64 (High Throughput)
Total Inference Memory

Required GPU VRAM

vLLM / SGLang
0.0 GB VRAM
Weights: 0 GB KV Cache: 0 GB Overhead: 2.5 GB
KV Cache per Stream: 0.0 GB Tokens Allocated: 0 tokens

GPU Hardware Compatibility Matrix

Which GPUs can host this configuration with zero Out-Of-Memory (OOM) errors.

✓ OOM Safety Buffer Added

The Mathematics of KV Cache Scaling

How Key-Value caching trades GPU high-bandwidth memory (HBM) for extreme inference throughput.

📐 Multi-Head Attention (MHA / GQA) Formula

Memory = 2 × L × H_kv × d_h × P × C × B
  • 2: Separate Key (K) and Value (V) tensors
  • L: Transformer layers (e.g., 32 for 8B, 80 for 70B)
  • H_kv: Number of Key-Value heads (8 in GQA)
  • d_h: Head dimension ($d_{\text{model}} / H_{\text{query}}$, typically 128)
  • P: Precision in bytes (2 for FP16, 1 for FP8, 0.5 for INT4)
  • C: Context length in tokens; B: Batch size

⚡ DeepSeek Multi-Head Latent Attention (MLA)

Memory = L × (d_c + d_R) × P × C × B

DeepSeek-V2 and V3 compress the Key-Value cache into a low-rank latent vector of dimension $d_c = 512$ plus decoupled RoPE dimension $d_R = 64$. This compresses KV cache memory by up to 93% compared to standard MHA, allowing 128k context on massive batch sizes.

Frequently Asked Questions

Essential engineering questions on KV cache allocation.

What is the formula for calculating LLM KV Cache size? ▼

For standard Multi-Head Attention (MHA) and Grouped-Query Attention (GQA), the formula is: KV Cache (bytes) = 2 × n_layers × n_kv_heads × d_head × precision_bytes × context_length × batch_size. The factor of 2 accounts for storing both Key and Value tensors. In FP16 precision, precision_bytes = 2; in FP8 it is 1, and in INT4 it is 0.5.

Why does the KV cache exceed model weights at large context windows? ▼

Model weights are static and do not scale with request length. However, the KV cache grows linearly with both context length and batch size. For example, running Llama-3-70B with FP16 KV cache across 32 concurrent streams at 64k context requires over 130 GB of VRAM just for the KV cache alone, vastly eclipsing the model weights.

How does Grouped-Query Attention (GQA) reduce KV cache memory? ▼

Standard Multi-Head Attention pairs every query head with its own key and value head. Grouped-Query Attention (GQA), used in models like Llama 3, Mistral, and Qwen, shares a single key-value head among multiple query heads (e.g., an 8:1 ratio in Llama 3 70B), reducing KV cache memory consumption by up to 8x with minimal impact on accuracy.

What is FP8 and INT4 KV cache quantization? ▼

Inference engines like vLLM, TensorRT-LLM, and SGLang support quantizing KV cache tensors to FP8 (8-bit floating point, cutting memory by 50%) or INT4 (4-bit integer, cutting memory by 75%). This allows higher batch sizes and longer context windows to fit onto smaller GPUs like the RTX 4090 or single A100.