GPU guide

Fix "CUDA out of memory" (and know when you need a bigger GPU)

Read the CUDA out of memory error, find what holds GPU memory, apply fixes in order of cost from batch size to a bigger GPU, and fix vLLM's startup errors.

Faiz Ahmed11 min read

The fix for CUDA out of memory starts with the error message, not with a bigger GPU. PyTorch prints how much the failed allocation asked for, how much the GPU has free, and how much your own process holds. Read those numbers, find what is holding the memory, then work down the fixes in order of what they cost you: batch size, sequence length, precision, gradient checkpointing, quantization, offload, and a GPU with more memory last. vLLM fails differently, and its startup errors have their own section below.

Read the error message#

Most of the answer is in the message itself. In PyTorch 2.14 it reads like this, wrapped here for reading and with placeholders where your sizes go:

Output
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate <request>. GPU 0 has a total
capacity of <total> of which <free> is free. Including non-PyTorch memory, this process has
<process> memory in use. Of the allocated memory <allocated> is allocated by PyTorch, and
<cached> is reserved by PyTorch but unallocated. If reserved but unallocated memory is large
try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.

Each number points somewhere different:

Part of the messageWhat it tells youWhere to look next
Tried to allocateThe one request that failedIf it is small and free memory is small too, the GPU was already full
total capacity, freeWhat the whole GPU has leftIf free is low but your process holds little, something else is using the GPU
this process has ... in useEverything your process holds, including memory outside PyTorchCompare it with the total capacity
Process N has ... in useAnother process on the same GPU, listed only when there is oneFind it with nvidia-smi and stop it
allocated by PyTorchLive tensors, the real size of your jobWork down the fixes below
reserved by PyTorch but unallocatedCached blocks that could not serve the requestIf it is large, fix fragmentation first

The exception class is torch.OutOfMemoryError, and torch.cuda.OutOfMemoryError names the same class, so either one works in an except clause.

Find what is holding the memory#

The one thing I always check first is whether my process is the only one on the GPU:

Terminal
nvidia-smi
nvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv

A notebook kernel you forgot, a second training run or an idle inference server all show up here with their PID. If memory stays in use after your script has exited, the PyTorch FAQ's answer is that a Python subprocess is usually still alive: find it with ps -elf | grep python and end it with kill -9 and its PID.

nvidia-smi alone can mislead, because PyTorch keeps memory it has freed in a cache for reuse, so the GPU looks fuller than your tensors need. Inside Python, ask PyTorch directly:

Python
import torch

free, total = torch.cuda.mem_get_info()
print(f"free {free / 2**30:.1f} of {total / 2**30:.1f} GiB")
print(f"allocated {torch.cuda.memory_allocated() / 2**30:.1f} GiB, "
      f"reserved {torch.cuda.memory_reserved() / 2**30:.1f} GiB, "
      f"peak {torch.cuda.max_memory_allocated() / 2**30:.1f} GiB")
print(torch.cuda.memory_summary())

For a job that fails partway through, record a memory snapshot and read it in PyTorch's viewer. Wrap the part that fails:

Python
torch.cuda.memory._record_memory_history()
try:
    trainer.train()  # or whatever loop runs out of memory
except torch.OutOfMemoryError:
    torch.cuda.memory._dump_snapshot("oom_snapshot.pickle")
    raise

Copy the file to your computer and drop it onto pytorch.org/memory_viz. The viewer runs in your browser and does not upload the snapshot. Its timeline shows the tensors that were alive over time, with the stack trace that allocated each one, so the allocation that tipped the GPU over is easy to find.

In Jupyter, memory often belongs to objects you no longer think about: a model from an earlier cell, or the exception itself, which the PyTorch FAQ notes keeps a reference to the frame where it was raised. Delete the references and empty the cache, or restart the kernel:

Python
import gc

del model, optimizer  # whatever still holds GPU tensors
gc.collect()
torch.cuda.empty_cache()

empty_cache() only hands cached blocks back. It cannot free memory that live tensors still hold, which is why the del comes first.

Two fixes that cost nothing#

Before you trade speed or quality for memory, rule out the two causes that are free to fix.

The first is holding references you do not need. The PyTorch FAQ's own example is a training loop that adds the loss tensor to a running total, which keeps every step's autograd graph alive. Add a plain number instead:

Python
total_loss += float(loss)  # not: total_loss += loss

For evaluation and generation, run the model under torch.inference_mode() so no graph is recorded at all. model.eval() does not do that on its own: it changes how layers such as dropout behave, and gradient tracking stays on.

Python
model.eval()
with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=256)

The second is fragmentation. If the message shows a large amount reserved by PyTorch but unallocated, the memory exists but in pieces too small for the request. Start Python with expandable segments:

Terminal
export PYTORCH_ALLOC_CONF=expandable_segments:True
python train.py

PYTORCH_ALLOC_CONF is the current name. The error message still suggests PYTORCH_CUDA_ALLOC_CONF, which PyTorch keeps as an alias, and the docs mark the option as experimental.

Fixes in order of cost#

Each step down this list costs more than the one before it, in speed, quality or money, so stop at the first one that works.

1. Lower the batch size and keep the effective batch

Activations scale with batch size, so a smaller per-device batch is the first real fix to try. Gradient accumulation keeps the effective batch, here 1 x 16 = 16, by running more, smaller steps before each optimizer update:

Python
args = SFTConfig(  # TrainingArguments takes the same two arguments
    per_device_train_batch_size = 1,
    gradient_accumulation_steps = 16,
    # ...your other settings
)

Unsloth gives the same advice for out-of-memory errors: set the batch size to 1, 2 or 3.

2. Shorten the sequence length

Activations also scale with sequence length, and so does the KV cache during inference. In TRL, max_length caps it (1,024 tokens by default in TRL 1.14), and in Unsloth it is max_seq_length. Check your data's length distribution before you cut, because truncation drops the end of every longer example. TRL's padding_free option, which works with FlashAttention 2 or 3, also saves the memory that padding tokens would take.

3. Use lower precision

Mixed-precision training with AdamW costs 18 bytes per parameter in Hugging Face's accounting: 6 for the weights, 8 for the optimizer states and 4 for the gradients. An 8-bit optimizer from bitsandbytes cuts the optimizer part from 8 bytes to 2. Run in bf16 mixed precision as well, which keeps the activations in 16-bit instead of 32-bit. TRL's docs ask for an Ampere or newer GPU for bf16, and every GPU in the QuantaCloud catalog qualifies, since the oldest, the RTX A6000 and A100, are Ampere.

Python
args = SFTConfig(bf16 = True, optim = "adamw_8bit")

4. Turn on gradient checkpointing

Gradient checkpointing keeps only some activations and recomputes the rest in the backward pass. The ZeRO paper's example is a 1.5B GPT-2 at sequence length 1K and batch size 32: about 60 GB of activations, and about 8 GB with checkpointing, for 33% more computation. Transformers puts the slowdown at about 20% and Axolotl at about 30%. TRL's SFTConfig turns it on by default as of 1.14, while plain TrainingArguments leaves it off:

Python
args = TrainingArguments(gradient_checkpointing = True)
# Unsloth: FastModel.get_peft_model(model, use_gradient_checkpointing = "unsloth", ...)

5. Quantize the base model

Loading a frozen base model in 4-bit is the QLoRA move: about half a byte per parameter instead of 2. With Transformers and bitsandbytes:

Python
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

bnb = BitsAndBytesConfig(
    load_in_4bit = True,
    bnb_4bit_quant_type = "nf4",
    bnb_4bit_compute_dtype = torch.bfloat16,
    bnb_4bit_use_double_quant = True,
)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config = bnb, dtype = "auto")

If you were running a full fine-tune, moving to LoRA or QLoRA is the bigger step, because gradients and optimizer states then exist only for a small adapter. The LoRA, QLoRA and full fine-tuning guide has the numbers by model size, and the Unsloth tutorial runs a QLoRA job end to end. For inference, load an FP8 or 4-bit checkpoint of the model instead of BF16.

6. Offload to CPU memory

Offload keeps part of the model or the optimizer in system RAM and moves it to the GPU when needed, which is slow and needs a lot of host memory. For inference with Transformers, device_map="auto" fills the GPUs first and puts the remaining layers in CPU RAM. For training, FSDP and DeepSpeed can offload parameters: Hugging Face PEFT's 70B QLoRA run dropped from 35.6 to 19.6 GB per GPU with offload, and used about 107 GB of host RAM to do it. Check the RAM of your configuration before you count on offload. In vLLM, --cpu-offload-gb does the same for weights.

7. Move to a GPU with more memory

When the job only fits with batch size 1, checkpointing on and a shortened sequence, and it still runs out, stop tuning and move up a tier. The usable column is our calculation at vLLM's default budget of 92%, which I also use as the working budget for training:

GPU memoryUsable at 92% (our calculation)QuantaCloud GPUs
48 GB41.4 GiB with ECC onRTX A6000, RTX 6000 Ada, L40, L40S
80 GB73.6 GiB on the A100, 73.3 GiB on the H100 PCIeA100 80GB, H100 PCIe
96 GB87.9 GiBRTX PRO 6000 Blackwell
141 GB129.2 GiBH200 NVL

Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes
H100 PCIe-Not listedNo
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 19:50 UTC

The rule I follow: if the weights plus the activations or KV cache of one full-size step do not fit in the usable column, move up a tier instead of fighting flags. Past 141 GB, split the job across 2, 4 or 8 GPUs in one VM.

Moving means a new instance, and stopping the old one deletes its disk. Push checkpoints and adapters to the Hugging Face Hub or object storage, or copy them off with scp, before you stop. The deploy docs cover launching the new one.

vLLM out of memory: the startup errors and their flags#

vLLM claims its memory at startup, so its out-of-memory errors usually arrive before the first request and name the setting to change. Version 0.30.0 has five, in the order they can appear:

Error begins withWhat it meansWhat to change
Free memory on device ... is less than desired GPU memory utilizationSomething else already holds GPU memory, often another server or a notebookStop the other process, or lower --gpu-memory-utilization
Failed to load modelThe weights alone do not fit on the GPUAn FP8 or 4-bit checkpoint, --tensor-parallel-size across more GPUs, or a bigger GPU
No available memory for the cache blocksThe weights and the startup profiling fill the whole budgetA higher --gpu-memory-utilization if nothing else shares the GPU, an FP8 or 4-bit checkpoint, --tensor-parallel-size across more GPUs, or a bigger GPU
To serve at least one request with the model's max seq lenThe KV cache for one full-length request does not fit--max-model-len at or below the length the message estimates, or --kv-cache-dtype fp8
CUDA out of memory occurred when warming up samplerThe warm-up batch of --max-num-seqs requests does not fit in what is leftLower --max-num-seqs or --gpu-memory-utilization

--gpu-memory-utilization defaults to 0.92 and applies per vLLM instance, so raising it helps only when nothing else shares the GPU. --max-model-len auto picks the longest context that fits. An FP8 KV cache halves the cache, but without calibration it uses scales of 1.0, so check the output quality on your own prompts. For the last few gigabytes, --max-num-seqs lowers how many requests vLLM batches at once, --enforce-eager gives up CUDA graphs and the memory they take at some cost in speed, and multimodal models can skip inputs you do not use with --limit-mm-per-prompt. A typical fix for a model that almost fits, kept on 127.0.0.1 because vLLM otherwise listens on the VM's public IP:

Terminal
vllm serve Qwen/Qwen3-8B \
  --host 127.0.0.1 \
  --max-model-len 16384 \
  --kv-cache-dtype fp8 \
  --max-num-seqs 64

After startup, vLLM logs a line that starts with GPU KV cache size: and gives the maximum concurrency at your --max-model-len, which tells you how much room the flags bought. When you split a model across GPUs without NVLink, such as the L40S, vLLM's docs recommend pipeline parallelism over tensor parallelism. The vLLM Docker guide and the vLLM install guide show these flags in full commands, and the KV cache guide sizes the cache for any model.


My order of operations: read the message, clear anything else off the GPU, then batch size, sequence length, precision and checkpointing. If the job still does not fit in the usable memory of the card, move up a tier rather than stacking workarounds, and use how much VRAM you need to pick it. The GPU catalog lists every tier with its live price.

Keep building

Choose your next step.