GPU guide

Fix vLLM out-of-memory errors

Fix vLLM out-of-memory errors: the five startup errors in vLLM 0.30.0, what each memory flag trades, and worked maths for 48 and 80 GB GPUs.

Faiz Ahmed12 min read

Most vLLM out-of-memory errors happen at startup, and the message names the setting to change. vLLM claims 92% of the GPU, loads the weights, measures what activations and CUDA graphs need, and gives the rest to the KV cache. If the weights plus the cache for one request at the full context length do not fit in that budget, it stops before serving anything. Fix it in this order: clear anything else off the GPU, lower --max-model-len, store the KV cache in FP8, lower --max-num-seqs, load a quantized checkpoint, split the model across GPUs, and only then rent a bigger GPU. A torch.OutOfMemoryError in your own training or inference code is a different problem, and fixing CUDA out of memory covers it.

How vLLM spends GPU memory#

The budget is simple arithmetic, and the startup log shows every term of it. vLLM requests --gpu-memory-utilization times the GPU's total memory, 0.92 by default. From that it subtracts the weights, the peak activation memory of a profiling run with a full step of dummy tokens, and, since vLLM 0.21, an estimate of the CUDA graph memory. What remains becomes the KV cache.

Output
KV cache = total memory x gpu_memory_utilization - weights - peak activations - CUDA graphs

Three log lines give you the terms, in this order:

Output
Model loading took <weights> GiB memory and <time> seconds
Available KV cache memory: <cache> GiB
GPU KV cache size: <tokens> tokens, Maximum concurrency for <max_model_len> tokens per request: <x>x

The last line is the one to read. The token count is the pool that every running request shares, and the concurrency figure is how many requests of the full --max-model-len fit in it at once. For Qwen3-8B in BF16 on a 48 GB RTX A6000 with ECC on, the arithmetic leaves about 26.1 GiB after the weights, room for about 190,000 tokens at 144 KiB each, or about 210,000 with ECC off (our calculation, before activations). The KV cache guide has the per-token formula for any model.

The five startup errors#

vLLM 0.30.0 has five out-of-memory errors at startup, and each one points at a different part of the budget. They appear in this order:

Error begins withWhat ran outWhat to change
Free memory on device ... is less than desired GPU memory utilizationThe GPU was not free when vLLM startedStop the other process, or lower --gpu-memory-utilization
Failed to load model - not enough GPU memoryThe weights aloneA quantized checkpoint, --tensor-parallel-size, or a bigger GPU
No available memory for the cache blocksNothing left after weights and activationsRaise --gpu-memory-utilization if the GPU is yours alone, or load smaller weights
To serve at least one request with the model's max seq lenThe KV cache for one full-length request--max-model-len at or below vLLM's estimate, or --kv-cache-dtype fp8
CUDA out of memory occurred when warming up samplerThe warm-up batch of --max-num-seqs requestsLower --max-num-seqs, or --gpu-memory-utilization

The first error is about other processes, and its full text gives you the numbers:

Output
Free memory on device <device> (<free>/<total> GiB) on startup is less than desired
GPU memory utilization (<utilization>, <requested> GiB). Decrease GPU memory
utilization or reduce GPU memory used by other processes.

Something already holds memory, for example a previous vLLM container, a notebook kernel or an Ollama server. Find it with nvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv and stop it. On a single-GPU VM that you rent for vLLM, lowering the fraction only hides the problem.

The fourth error tells you the fix in its own text:

Output
To serve at least one request with the model's max seq len (<max_model_len>), (<needed> GiB
KV cache is needed, which is larger than the available KV cache memory (<available> GiB).
Based on the available memory, the estimated maximum model length is <estimate>. Try
increasing `gpu_memory_utilization` (which also controls CPU memory on the CPU backend)
or decreasing `max_model_len` when initializing the engine.

Set --max-model-len at or below the estimated length and start again, or pass --max-model-len auto and vLLM picks the largest length that fits. The Failed to load model message also suggests lowering --gpu-memory-utilization, but the weights do not shrink with the budget: when they alone do not fit, the fix is smaller weights, more GPUs or a bigger GPU.

--gpu-memory-utilization#

Raise it only when nothing else uses the GPU. The default of 0.92 applies per vLLM instance, and vLLM's docs give the example of two instances on one GPU at 0.5 each. Going from 0.92 to 0.95 on a 48 GB card adds about 1.4 GiB to the KV cache (our calculation: 0.03 x 45 GiB), which is worth trying when the fourth error misses by a little. Going higher leaves less room for memory outside vLLM's accounting, and vLLM's own docstring warns that too high a value can cause out-of-memory errors.

If you would rather set the cache in bytes, --kv-cache-memory-bytes sizes it directly and ignores --gpu-memory-utilization.

--max-model-len#

The context limit is the biggest lever you have. vLLM sizes the cache so that at least one request of --max-model-len tokens fits, and by default it takes that length from the model's config. For Llama 3.3 70B that is 131,072 tokens, and at 320 KiB per token one such request needs 40 GiB of KV cache on its own (our calculation). If your prompts and answers never exceed 16,384 tokens, say so: the same request budget drops to 5 GiB.

Write the number out in full. vLLM reads 32k as 32,000 tokens and 32K as 32,768. A request longer than the limit is rejected with This model's maximum context length is ... tokens, not truncated, so pick a limit that covers your longest real prompt plus its answer.

--max-num-seqs and --max-num-batched-tokens#

These two cap how much work runs at once, and with it the activation memory. --max-num-seqs is the number of requests in a step: 256 by default on 48 GB cards and the A100, and 1,024 on GPUs with 70 GiB or more such as the H100, H200 NVL and RTX PRO 6000. vLLM warms up its sampler with that many dummy requests, which is where the fifth error comes from. Its message suggests a lower --max-num-seqs or --gpu-memory-utilization, and I would try 64 or 128 sequences first, because a lower utilization also shrinks the KV cache. --max-num-batched-tokens caps the tokens per step, 2,048 by default for the server on 48 GB cards and the A100, and 8,192 on the larger GPUs.

The same limits govern the out-of-memory problem that does not crash. When running requests fill the KV cache, vLLM preempts some of them and recomputes them later, and the periodic stats line gains a Preemptions: count next to GPU KV cache usage:. vLLM's tuning guide lists the remedies: a higher --gpu-memory-utilization, a lower --max-num-seqs or --max-num-batched-tokens, or more GPUs through tensor or pipeline parallelism.

Store the KV cache in FP8#

--kv-cache-dtype fp8 halves the cache on any GPU, which doubles the tokens that fit. It works independently of the weight format. Without a calibrated checkpoint vLLM uses scales of 1.0, so compare outputs on your own prompts before you rely on it.

Load a quantized checkpoint#

Smaller weights leave more for everything else. Llama 3.3 70B is 141.1 GB in BF16, 72.7 GB as per-channel FP8 and 39.8 GB as 4-bit AWQ, so an 80 GB card goes from not loading the model at BF16, to holding it with little room for context at FP8, to holding it with room to spare at 4 bits. Which format runs at full speed depends on the GPU: FP8 math needs Ada or newer, and on the A100 and RTX A6000 FP8 only saves memory. The vLLM quantization guide matches formats to GPUs.

Split the model across GPUs#

Tensor parallelism divides the weights and the KV cache across the GPUs of one VM, and each GPU gets its own 92% budget. On a 2x VM, add --tensor-parallel-size 2. vLLM's docs recommend pipeline parallelism (--pipeline-parallel-size 2) on GPUs without NVLink, such as the L40S, and QuantaCloud offers multi-GPU VMs with up to 8 GPUs. vLLM across several GPUs covers the choice between the two.

Offload and eager mode#

Two more flags buy the last few gigabytes at a cost in speed. --cpu-offload-gb 10 keeps 10 GiB of weights in system RAM and moves them to the GPU on every forward pass, which slows every step. --enforce-eager skips CUDA graph capture and gives back the memory the graphs take, at some cost in decode speed. For a multimodal model, --limit-mm-per-prompt stops vLLM from reserving memory for images or video you never send. Ollama handles a model that does not fit differently, by keeping part of it in system RAM on its own, which vLLM vs Ollama covers.

Worked example: a 70B model on a 48 GB GPU#

The 4-bit AWQ build of Llama 3.3 70B loads on a 48 GB card, but not at its default context. The numbers, as our calculation from the 46,068 MiB such a card reports with ECC on:

StepAmount
vLLM's budget, 0.92 x 45.0 GiB41.4 GiB
Weights, casperhansen/llama-3.3-70b-instruct-awq (39.8 GB)37.0 GiB
Left for KV cache, before activations and CUDA graphs4.4 GiB
One request at the default 131,072 tokens40 GiB
Tokens that fit in 4.4 GiB, BF16 cacheabout 14,000
Tokens that fit in 4.4 GiB, FP8 cacheabout 28,000

At the default length vLLM stops with the max seq len error. A limit of 16,384 needs 5 GiB, over that upper bound with a BF16 cache, while 8,192 needs 2.5 GiB, and --max-model-len auto finds the longest length that really fits once activations and CUDA graphs have their share. An RTX A6000 or RTX 6000 Ada with ECC off reports 49,140 MiB, which adds 2.8 GiB to every figure above. The command below uses vllm serve from installing vLLM, and with Docker the same flags go after the image name, as in the vLLM Docker guide:

Terminal
vllm serve casperhansen/llama-3.3-70b-instruct-awq \
  --host 127.0.0.1 \
  --max-model-len auto \
  --api-key "$VLLM_API_KEY"

A 48 GB card makes that model a one-conversation server, and several users at once need the 80 GB card in the next section.

Worked example: a 70B model on an 80 GB GPU#

On an 80 GB card the choice between FP8 and 4-bit decides everything. The FP8 checkpoint, 72.7 GB or 67.7 GiB, leaves 5.9 GiB of the A100's 73.6 GiB budget, about 19,000 tokens at BF16 or 38,000 with an FP8 cache, and 5.6 GiB, about 18,000 tokens, of the H100 PCIe's 73.3 GiB. On the A100 those FP8 weights also run weight-only, because Ampere has no FP8 math, while the H100 PCIe runs them natively. The AWQ checkpoint leaves 36.6 GiB on the A100, about 120,000 tokens, and 36.2 GiB, about 119,000, on the H100 PCIe: room for several long conversations at once, though not for one at the full 131,072-token context, which needs 40 GiB (our calculations).

When to move up a size#

The rule I follow: when the weights plus the KV cache for the context and concurrency you need do not fit in 92% of the card, move up a tier rather than stacking flags. Each flag above trades away context, speed or quality, and I would rather pay for the next tier than run a server that holds one short conversation.

GPU memoryvLLM budget at 0.92 (our calculation)QuantaCloud GPUs
48 GB41.4 GiB with ECC onRTX A6000, RTX 6000 Ada, L40, L40S
80 GB73.6 GiB on the A100, 73.3 GiB on the H100 PCIeA100 80GB, H100 PCIe
96 GB87.9 GiBRTX PRO 6000 Blackwell
141 GB129.2 GiBH200 NVL
GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 20:56 UTC

Moving means a new instance. Stopping the old one terminates it and deletes its disk, so the new one downloads the image and the weights again, and anything you want to keep has to be copied off first.

vLLM out-of-memory FAQ#

What should --gpu-memory-utilization be?

Leave it at 0.92 when vLLM has the GPU to itself, and lower it only to share the GPU with another process. Raising it to 0.95 buys a little more KV cache when the max seq len error misses by a small margin.

Why does nvidia-smi show the GPU almost full with no requests running?

Because vLLM claims its whole budget at startup and turns what is left after the weights into KV cache blocks. The memory is reserved, not leaked, and the GPU KV cache usage figure in vLLM's stats line shows how much of it is in use.

Does a lower --max-model-len truncate my prompts?

No. vLLM rejects a request whose prompt plus requested output exceeds the limit, with an error that starts This model's maximum context length is. Set the limit to cover your longest real request.

Is --kv-cache-dtype fp8 safe to use?

It halves the cache at no cost in memory elsewhere, which makes it the next fix I try after --max-model-len. Without calibration vLLM uses scales of 1.0, so compare answers on your own prompts before you rely on it.

What happens when the KV cache fills after startup?

vLLM preempts some running requests and recomputes them when space frees, rather than failing them, and the stats line shows a Preemptions: count. If that count keeps rising, lower --max-num-seqs, give vLLM more memory, or add GPUs.


My order of operations: read which of the five errors you have, clear the GPU, set --max-model-len to what you really need, then try an FP8 cache and a smaller --max-num-seqs. If a model only fits with a context too short for your users, stop tuning and move up a tier, and use how much VRAM you need to pick it. The GPU catalog lists every tier with its live price.

Launch an H200 NVL (141 GB) Launch an RTX PRO 6000 (96 GB)

Keep building

Choose your next step.