The KV cache is the memory an LLM server uses to remember every token it has already processed. For each token in each active request, every attention layer stores a key vector and a value vector, so the model does not recompute them when it generates the next token. It grows linearly with the context length and with the number of requests running at once, and on a serving GPU it is usually what decides how many users and how much context you can have. The weights decide whether a model loads. The KV cache decides whether it is useful.
What the KV cache stores#
The cache exists because of how generation works. In prefill, the model reads the whole prompt in one pass and writes a key and a value for every prompt token into every layer. In decode, each new token has to attend to all the tokens before it, so the model reads their keys and values back, computes the new token's pair and appends it. Without the cache, every new token would recompute the entire history.
The numbers get large quickly. The paper that introduced vLLM works one out for OPT-13B: 800 KB per token (2 x 5,120 hidden size x 40 layers x 2 bytes), so a single 2,048-token request can hold 1.6 GB. Serving a 13B model on a 40 GB A100, the weights took about 65% of the memory and the KV cache close to 30%. That 30% is what limits how many requests the server can batch together.
The KV cache formula#
The formula is one line:
KV cache bytes = 2 x layers x KV heads x head_dim x bytes per value x tokens x concurrent requests
The 2 counts keys and values. The other terms come from the model's config.json. Layers is num_hidden_layers, counting only the layers whose cache grows with context. KV heads is num_key_value_heads, not num_attention_heads. head_dim is in the config, or it is hidden_size divided by num_attention_heads when the config has none. Bytes per value is 2 for BF16 or FP16 and 1 for an FP8 cache.
The easy mistake is to use the query-head count or hidden_size, as the total-size formula in NVIDIA's inference optimization guide does. That formula assumes every head keeps its own keys and values, which is not true of any model in the table below. For Qwen3-8B the correct figure is 2 x 36 x 8 x 128 x 2 = 147,456 bytes, or 144 KiB per token. The hidden_size version gives 576 KiB, four times too high (our calculation).
This script, saved as kv_cache.py, is a KV cache size calculator for any Hugging Face model. It reads the config, handles standard, grouped-query and MLA attention, and skips sliding-window and linear-attention layers:
import json, sys
from huggingface_hub import hf_hub_download
repo, tokens = sys.argv[1], int(sys.argv[2])
cfg = json.load(open(hf_hub_download(repo, "config.json")))
cfg = cfg.get("text_config", cfg)
if "kv_lora_rank" in cfg: # MLA: one compressed latent per layer
per_token = cfg["num_hidden_layers"] * (cfg["kv_lora_rank"] + cfg["qk_rope_head_dim"]) * 2
else: # MHA or GQA; sliding-window and linear-attention layers are skipped
types = cfg.get("layer_types")
layers = types.count("full_attention") if types else cfg["num_hidden_layers"]
kv_heads = cfg.get("num_key_value_heads") or cfg["num_attention_heads"]
head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
per_token = 2 * layers * kv_heads * head_dim * 2 # K and V at 2 bytes (BF16)
print(f"{repo}: {per_token:,} bytes per token, "
f"{per_token * tokens / 2**30:.2f} GiB for {tokens:,} tokens")
$ python kv_cache.py Qwen/Qwen3-8B 32768
Qwen/Qwen3-8B: 147,456 bytes per token, 4.50 GiB for 32,768 tokens
$ python kv_cache.py openai/gpt-oss-120b 131072
openai/gpt-oss-120b: 36,864 bytes per token, 4.50 GiB for 131,072 tokens
It needs only huggingface_hub, which every vLLM environment already has. Set HF_TOKEN for gated repos such as Llama. It overstates Gemma 4, whose global layers use their own head settings, and it ignores small fixed-size extras such as a linear-attention state.
KV cache sizes for popular models#
Every row below comes from the model's own config (our calculation, BF16 cache):
| Model | Attention | Layers that grow the cache | KV heads x head_dim | KV per token | 32,768 tokens | Full context |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | GQA | 32 | 8 x 128 | 128 KiB | 4 GiB | 16 GiB at 131,072 |
| Qwen3-8B | GQA | 36 | 8 x 128 | 144 KiB | 4.5 GiB | 5.6 GiB at 40,960 |
| Qwen3-32B | GQA | 64 | 8 x 128 | 256 KiB | 8 GiB | 10 GiB at 40,960 |
| Llama 3.3 70B | GQA | 80 | 8 x 128 | 320 KiB | 10 GiB | 40 GiB at 131,072 |
| gpt-oss-120b | GQA, 128-token window on half the layers | 18 of 36 | 8 x 64 | 36 KiB | 1.1 GiB | 4.5 GiB at 131,072 |
| Qwen3.8-27B | GQA plus linear attention | 16 of 64 | 4 x 256 | 64 KiB | 2 GiB | 16 GiB at 262,144 |
| Kimi-K2.6 | MLA | 61 | one latent of 512 + 64 | 68.6 KiB | 2.1 GiB | 17.2 GiB at 262,144 |
Two rows stand out. Llama 3.3 70B needs 40 GiB (42.9 GB) of cache for one full-length conversation, more than its 4-bit AWQ checkpoint takes (39.8 GB). And gpt-oss-120b, a 117B-parameter model, stores less per token than Qwen3-8B, because only half of its layers keep a growing cache and each of those has 64-dimension heads.
How GQA, MLA and sliding windows shrink it#
Four designs shrink the KV cache, and the models in that table use all of them.
Grouped-query attention lets several query heads share one key-value head. Llama 3.3 70B has 64 query heads and 8 KV heads, so its cache is one eighth of what full multi-head attention would need: 320 KiB per token instead of 2.5 MiB, and 40 GiB for a 131,072-token conversation instead of 320 GiB (our calculation). Multi-query attention takes the idea to its limit, with a single KV head.
Multi-head latent attention (MLA) caches one compressed latent per layer, plus a small positional key, instead of per-head keys and values. The DeepSeek-V2 paper puts its cache at the equivalent of grouped-query attention with only 2.25 groups, and reports a 93.3% smaller KV cache than DeepSeek 67B. That is how Kimi-K2.6, a model with about a trillion parameters, gets by with 70,272 bytes per token.
Sliding-window attention limits a layer to the most recent tokens, so that layer never caches more than its window. The Mistral 7B paper reports 8 times less cache memory at 32k tokens with its 4,096-token window. gpt-oss uses a 128-token window on every other layer, and Gemma-style models mix local and global layers the same way.
Linear-attention layers, such as the Gated DeltaNet layers in Qwen3.8-27B, keep a fixed-size state instead of a cache that grows. 48 of that model's 64 layers work this way, so only 16 layers grow with context.
On top of any of these, an FP8 cache halves the bytes per value. In vLLM that is --kv-cache-dtype fp8, which uses scales of 1.0 unless the checkpoint was calibrated, so check output quality on your own prompts.
How vLLM manages the KV cache#
vLLM gives the KV cache whatever memory is left. At startup it claims --gpu-memory-utilization of the GPU (0.92 by default), loads the weights, measures the peak activation memory, and turns the rest into fixed-size KV blocks. PagedAttention allocates those blocks as each request grows instead of reserving its maximum length up front. The vLLM paper found that earlier serving systems used only 20.4% to 38.2% of their KV cache memory for actual tokens, against 96.3% in vLLM.
Once the model is loaded, the log tells you the budget:
GPU KV cache size: N tokens, Maximum concurrency for 32,768 tokens per request: X.XXx
The token figure is shared by every running request. The concurrency figure is how many requests of the full --max-model-len fit at once. If not even one fits, vLLM refuses to start with an error that begins To serve at least one request with the model's max seq len and names the longest length that would fit. Lower --max-model-len or pass --max-model-len auto and let vLLM choose. Prefix caching is on by default, so requests that share a system prompt or a document share its blocks too.
The formula gives an upper bound, and the log gives the real number. For Qwen3-8B on an RTX A6000 at the default 0.92, the formula allows at most about 190,000 tokens (26.1 GiB after the weights, at 144 KiB each), or about 210,000 if the card runs with ECC off. For gpt-oss-120b on an H200 NVL, the formula allows about 2 million tokens.
The gap is what the formula leaves out: activation memory and CUDA graphs, which vLLM sets aside first, and any difference between the memory the formula assumes and the memory the card reports.
How the KV cache picks your GPU#
The rule I follow is to multiply the cache per token by the context you actually serve and by the number of requests you expect at once, add the weights, and pick the smallest card where the total stays under 92%. This table shows what is left for the cache after the weights, taken as the size of each checkpoint's files, and how many tokens that holds (our calculation from the memory each card reports in nvidia-smi, with ECC on for the 48 GB cards, before activations):
| Workload (weights) | 48 GB: RTX A6000, L40S | 80 GB: A100 | 80 GB: H100 PCIe | 96 GB: RTX PRO 6000 | 141 GB: H200 NVL |
|---|---|---|---|---|---|
| Qwen3-8B BF16 (15.3 GiB) | 26.1 GiB, about 190k tokens | 58.3 GiB, about 425k | 58.0 GiB, about 420k | 72.7 GiB, about 530k | 113.9 GiB, about 830k |
| Qwen3-32B FP8 (32.0 GiB) | 9.4 GiB, about 38k tokens | 41.6 GiB, about 170k | 41.3 GiB, about 170k | 56.0 GiB, about 230k | 97.2 GiB, about 400k |
| Llama 3.3 70B FP8 (67.7 GiB) | Does not fit | 5.9 GiB, about 19k | 5.6 GiB, about 18k | 20.3 GiB, about 66k | 61.5 GiB, about 200k |
| gpt-oss-120b MXFP4 (60.8 GiB) | Does not fit | 12.8 GiB, about 370k | 12.5 GiB, about 360k | 27.2 GiB, about 790k | 68.4 GiB, about 2M |
Read each cell as the pool that all active requests share. Llama 3.3 70B at FP8 on an 80 GB card has room for 18,000 to 19,000 tokens: one long conversation, or a handful of short ones. On an H200 NVL the same model holds about 200,000, enough for six 32k conversations at once. That difference, not the weights, is the case for the 141 GB card. An 8B model is the opposite case: even a 48 GB RTX A6000 or L40S holds five full 32k conversations, and a bigger card mainly buys concurrency. A 32B model at FP8 is the borderline case on 48 GB: one 32k conversation needs 8 GiB of the 9.4 GiB left, before activations. With ECC off, an RTX A6000 reports 49,140 MiB and each 48 GB cell gains 2.8 GiB.
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
The A100, H100 and RTX PRO 6000 pages list what else fits in their memory, and how much VRAM you need covers weights for other workloads. If you already hit the limit, fixing CUDA out of memory goes through the fixes in order, and the gpt-oss sizing guide works the same numbers for one model family.
My rule is to work out the KV cache before choosing a GPU, not after the first out-of-memory error. Take the per-token figure from the script, multiply it by your context and your concurrency, add the weights, and pick the smallest card where that stays under 92%. Then confirm it on the first run with the KV cache line in vLLM's log, using the vLLM Docker guide.
Launch a GPU and check your KV cache