GPU guide

KV cache explained: how much GPU memory LLM serving really needs

What the KV cache is, the formula to size it, worked numbers for Llama, Qwen, gpt-oss and Kimi, and how it decides between 48, 80, 96 and 141 GB GPUs.

Faiz Ahmed10 min read

The KV cache is the memory an LLM server uses to remember every token it has already processed. For each token in each active request, every attention layer stores a key vector and a value vector, so the model does not recompute them when it generates the next token. It grows linearly with the context length and with the number of requests running at once, and on a serving GPU it is usually what decides how many users and how much context you can have. The weights decide whether a model loads. The KV cache decides whether it is useful.

What the KV cache stores#

The cache exists because of how generation works. In prefill, the model reads the whole prompt in one pass and writes a key and a value for every prompt token into every layer. In decode, each new token has to attend to all the tokens before it, so the model reads their keys and values back, computes the new token's pair and appends it. Without the cache, every new token would recompute the entire history.

The numbers get large quickly. The paper that introduced vLLM works one out for OPT-13B: 800 KB per token (2 x 5,120 hidden size x 40 layers x 2 bytes), so a single 2,048-token request can hold 1.6 GB. Serving a 13B model on a 40 GB A100, the weights took about 65% of the memory and the KV cache close to 30%. That 30% is what limits how many requests the server can batch together.

The KV cache formula#

The formula is one line:

Output
KV cache bytes = 2 x layers x KV heads x head_dim x bytes per value x tokens x concurrent requests

The 2 counts keys and values. The other terms come from the model's config.json. Layers is num_hidden_layers, counting only the layers whose cache grows with context. KV heads is num_key_value_heads, not num_attention_heads. head_dim is in the config, or it is hidden_size divided by num_attention_heads when the config has none. Bytes per value is 2 for BF16 or FP16 and 1 for an FP8 cache.

The easy mistake is to use the query-head count or hidden_size, as the total-size formula in NVIDIA's inference optimization guide does. That formula assumes every head keeps its own keys and values, which is not true of any model in the table below. For Qwen3-8B the correct figure is 2 x 36 x 8 x 128 x 2 = 147,456 bytes, or 144 KiB per token. The hidden_size version gives 576 KiB, four times too high (our calculation).

This script, saved as kv_cache.py, is a KV cache size calculator for any Hugging Face model. It reads the config, handles standard, grouped-query and MLA attention, and skips sliding-window and linear-attention layers:

Python
import json, sys
from huggingface_hub import hf_hub_download

repo, tokens = sys.argv[1], int(sys.argv[2])
cfg = json.load(open(hf_hub_download(repo, "config.json")))
cfg = cfg.get("text_config", cfg)
if "kv_lora_rank" in cfg:  # MLA: one compressed latent per layer
    per_token = cfg["num_hidden_layers"] * (cfg["kv_lora_rank"] + cfg["qk_rope_head_dim"]) * 2
else:  # MHA or GQA; sliding-window and linear-attention layers are skipped
    types = cfg.get("layer_types")
    layers = types.count("full_attention") if types else cfg["num_hidden_layers"]
    kv_heads = cfg.get("num_key_value_heads") or cfg["num_attention_heads"]
    head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
    per_token = 2 * layers * kv_heads * head_dim * 2  # K and V at 2 bytes (BF16)
print(f"{repo}: {per_token:,} bytes per token, "
      f"{per_token * tokens / 2**30:.2f} GiB for {tokens:,} tokens")
Output
$ python kv_cache.py Qwen/Qwen3-8B 32768
Qwen/Qwen3-8B: 147,456 bytes per token, 4.50 GiB for 32,768 tokens
$ python kv_cache.py openai/gpt-oss-120b 131072
openai/gpt-oss-120b: 36,864 bytes per token, 4.50 GiB for 131,072 tokens

It needs only huggingface_hub, which every vLLM environment already has. Set HF_TOKEN for gated repos such as Llama. It overstates Gemma 4, whose global layers use their own head settings, and it ignores small fixed-size extras such as a linear-attention state.

Every row below comes from the model's own config (our calculation, BF16 cache):

ModelAttentionLayers that grow the cacheKV heads x head_dimKV per token32,768 tokensFull context
Llama 3.1 8BGQA328 x 128128 KiB4 GiB16 GiB at 131,072
Qwen3-8BGQA368 x 128144 KiB4.5 GiB5.6 GiB at 40,960
Qwen3-32BGQA648 x 128256 KiB8 GiB10 GiB at 40,960
Llama 3.3 70BGQA808 x 128320 KiB10 GiB40 GiB at 131,072
gpt-oss-120bGQA, 128-token window on half the layers18 of 368 x 6436 KiB1.1 GiB4.5 GiB at 131,072
Qwen3.8-27BGQA plus linear attention16 of 644 x 25664 KiB2 GiB16 GiB at 262,144
Kimi-K2.6MLA61one latent of 512 + 6468.6 KiB2.1 GiB17.2 GiB at 262,144

Two rows stand out. Llama 3.3 70B needs 40 GiB (42.9 GB) of cache for one full-length conversation, more than its 4-bit AWQ checkpoint takes (39.8 GB). And gpt-oss-120b, a 117B-parameter model, stores less per token than Qwen3-8B, because only half of its layers keep a growing cache and each of those has 64-dimension heads.

How GQA, MLA and sliding windows shrink it#

Four designs shrink the KV cache, and the models in that table use all of them.

Grouped-query attention lets several query heads share one key-value head. Llama 3.3 70B has 64 query heads and 8 KV heads, so its cache is one eighth of what full multi-head attention would need: 320 KiB per token instead of 2.5 MiB, and 40 GiB for a 131,072-token conversation instead of 320 GiB (our calculation). Multi-query attention takes the idea to its limit, with a single KV head.

Multi-head latent attention (MLA) caches one compressed latent per layer, plus a small positional key, instead of per-head keys and values. The DeepSeek-V2 paper puts its cache at the equivalent of grouped-query attention with only 2.25 groups, and reports a 93.3% smaller KV cache than DeepSeek 67B. That is how Kimi-K2.6, a model with about a trillion parameters, gets by with 70,272 bytes per token.

Sliding-window attention limits a layer to the most recent tokens, so that layer never caches more than its window. The Mistral 7B paper reports 8 times less cache memory at 32k tokens with its 4,096-token window. gpt-oss uses a 128-token window on every other layer, and Gemma-style models mix local and global layers the same way.

Linear-attention layers, such as the Gated DeltaNet layers in Qwen3.8-27B, keep a fixed-size state instead of a cache that grows. 48 of that model's 64 layers work this way, so only 16 layers grow with context.

On top of any of these, an FP8 cache halves the bytes per value. In vLLM that is --kv-cache-dtype fp8, which uses scales of 1.0 unless the checkpoint was calibrated, so check output quality on your own prompts.

How vLLM manages the KV cache#

vLLM gives the KV cache whatever memory is left. At startup it claims --gpu-memory-utilization of the GPU (0.92 by default), loads the weights, measures the peak activation memory, and turns the rest into fixed-size KV blocks. PagedAttention allocates those blocks as each request grows instead of reserving its maximum length up front. The vLLM paper found that earlier serving systems used only 20.4% to 38.2% of their KV cache memory for actual tokens, against 96.3% in vLLM.

Once the model is loaded, the log tells you the budget:

Output
GPU KV cache size: N tokens, Maximum concurrency for 32,768 tokens per request: X.XXx

The token figure is shared by every running request. The concurrency figure is how many requests of the full --max-model-len fit at once. If not even one fits, vLLM refuses to start with an error that begins To serve at least one request with the model's max seq len and names the longest length that would fit. Lower --max-model-len or pass --max-model-len auto and let vLLM choose. Prefix caching is on by default, so requests that share a system prompt or a document share its blocks too.

The formula gives an upper bound, and the log gives the real number. For Qwen3-8B on an RTX A6000 at the default 0.92, the formula allows at most about 190,000 tokens (26.1 GiB after the weights, at 144 KiB each), or about 210,000 if the card runs with ECC off. For gpt-oss-120b on an H200 NVL, the formula allows about 2 million tokens.

The gap is what the formula leaves out: activation memory and CUDA graphs, which vLLM sets aside first, and any difference between the memory the formula assumes and the memory the card reports.

How the KV cache picks your GPU#

The rule I follow is to multiply the cache per token by the context you actually serve and by the number of requests you expect at once, add the weights, and pick the smallest card where the total stays under 92%. This table shows what is left for the cache after the weights, taken as the size of each checkpoint's files, and how many tokens that holds (our calculation from the memory each card reports in nvidia-smi, with ECC on for the 48 GB cards, before activations):

Workload (weights)48 GB: RTX A6000, L40S80 GB: A10080 GB: H100 PCIe96 GB: RTX PRO 6000141 GB: H200 NVL
Qwen3-8B BF16 (15.3 GiB)26.1 GiB, about 190k tokens58.3 GiB, about 425k58.0 GiB, about 420k72.7 GiB, about 530k113.9 GiB, about 830k
Qwen3-32B FP8 (32.0 GiB)9.4 GiB, about 38k tokens41.6 GiB, about 170k41.3 GiB, about 170k56.0 GiB, about 230k97.2 GiB, about 400k
Llama 3.3 70B FP8 (67.7 GiB)Does not fit5.9 GiB, about 19k5.6 GiB, about 18k20.3 GiB, about 66k61.5 GiB, about 200k
gpt-oss-120b MXFP4 (60.8 GiB)Does not fit12.8 GiB, about 370k12.5 GiB, about 360k27.2 GiB, about 790k68.4 GiB, about 2M

Read each cell as the pool that all active requests share. Llama 3.3 70B at FP8 on an 80 GB card has room for 18,000 to 19,000 tokens: one long conversation, or a handful of short ones. On an H200 NVL the same model holds about 200,000, enough for six 32k conversations at once. That difference, not the weights, is the case for the 141 GB card. An 8B model is the opposite case: even a 48 GB RTX A6000 or L40S holds five full 32k conversations, and a bigger card mainly buys concurrency. A 32B model at FP8 is the borderline case on 48 GB: one 32k conversation needs 8 GiB of the 9.4 GiB left, before activations. With ECC off, an RTX A6000 reports 49,140 MiB and each 48 GB cell gains 2.8 GiB.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

The A100, H100 and RTX PRO 6000 pages list what else fits in their memory, and how much VRAM you need covers weights for other workloads. If you already hit the limit, fixing CUDA out of memory goes through the fixes in order, and the gpt-oss sizing guide works the same numbers for one model family.


My rule is to work out the KV cache before choosing a GPU, not after the first out-of-memory error. Take the per-token figure from the script, multiply it by your context and your concurrency, add the weights, and pick the smallest card where that stays under 92%. Then confirm it on the first run with the KV cache line in vLLM's log, using the vLLM Docker guide.

Launch a GPU and check your KV cache

Keep building

Choose your next step.