The fix for CUDA out of memory starts with the error message, not with a bigger GPU. PyTorch prints how much the failed allocation asked for, how much the GPU has free, and how much your own process holds. Read those numbers, find what is holding the memory, then work down the fixes in order of what they cost you: batch size, sequence length, precision, gradient checkpointing, quantization, offload, and a GPU with more memory last. vLLM fails differently, and its startup errors have their own section below.
Read the error message#
Most of the answer is in the message itself. In PyTorch 2.14 it reads like this, wrapped here for reading and with placeholders where your sizes go:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate <request>. GPU 0 has a total
capacity of <total> of which <free> is free. Including non-PyTorch memory, this process has
<process> memory in use. Of the allocated memory <allocated> is allocated by PyTorch, and
<cached> is reserved by PyTorch but unallocated. If reserved but unallocated memory is large
try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.
Each number points somewhere different:
| Part of the message | What it tells you | Where to look next |
|---|---|---|
| Tried to allocate | The one request that failed | If it is small and free memory is small too, the GPU was already full |
| total capacity, free | What the whole GPU has left | If free is low but your process holds little, something else is using the GPU |
| this process has ... in use | Everything your process holds, including memory outside PyTorch | Compare it with the total capacity |
| Process N has ... in use | Another process on the same GPU, listed only when there is one | Find it with nvidia-smi and stop it |
| allocated by PyTorch | Live tensors, the real size of your job | Work down the fixes below |
| reserved by PyTorch but unallocated | Cached blocks that could not serve the request | If it is large, fix fragmentation first |
The exception class is torch.OutOfMemoryError, and torch.cuda.OutOfMemoryError names the same class, so either
one works in an except clause.
Find what is holding the memory#
The one thing I always check first is whether my process is the only one on the GPU:
nvidia-smi
nvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv
A notebook kernel you forgot, a second training run or an idle inference server all show up here with their PID. If
memory stays in use after your script has exited, the PyTorch FAQ's answer is that a Python subprocess is usually
still alive: find it with ps -elf | grep python and end it with kill -9 and its PID.
nvidia-smi alone can mislead, because PyTorch keeps memory it has freed in a cache for reuse, so the GPU looks
fuller than your tensors need. Inside Python, ask PyTorch directly:
import torch
free, total = torch.cuda.mem_get_info()
print(f"free {free / 2**30:.1f} of {total / 2**30:.1f} GiB")
print(f"allocated {torch.cuda.memory_allocated() / 2**30:.1f} GiB, "
f"reserved {torch.cuda.memory_reserved() / 2**30:.1f} GiB, "
f"peak {torch.cuda.max_memory_allocated() / 2**30:.1f} GiB")
print(torch.cuda.memory_summary())
For a job that fails partway through, record a memory snapshot and read it in PyTorch's viewer. Wrap the part that fails:
torch.cuda.memory._record_memory_history()
try:
trainer.train() # or whatever loop runs out of memory
except torch.OutOfMemoryError:
torch.cuda.memory._dump_snapshot("oom_snapshot.pickle")
raise
Copy the file to your computer and drop it onto pytorch.org/memory_viz. The viewer runs in your browser and does not upload the snapshot. Its timeline shows the tensors that were alive over time, with the stack trace that allocated each one, so the allocation that tipped the GPU over is easy to find.
In Jupyter, memory often belongs to objects you no longer think about: a model from an earlier cell, or the exception itself, which the PyTorch FAQ notes keeps a reference to the frame where it was raised. Delete the references and empty the cache, or restart the kernel:
import gc
del model, optimizer # whatever still holds GPU tensors
gc.collect()
torch.cuda.empty_cache()
empty_cache() only hands cached blocks back. It cannot free memory that live tensors still hold, which is why the
del comes first.
Two fixes that cost nothing#
Before you trade speed or quality for memory, rule out the two causes that are free to fix.
The first is holding references you do not need. The PyTorch FAQ's own example is a training loop that adds the loss tensor to a running total, which keeps every step's autograd graph alive. Add a plain number instead:
total_loss += float(loss) # not: total_loss += loss
For evaluation and generation, run the model under torch.inference_mode() so no graph is recorded at all.
model.eval() does not do that on its own: it changes how layers such as dropout behave, and gradient tracking stays
on.
model.eval()
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=256)
The second is fragmentation. If the message shows a large amount reserved by PyTorch but unallocated, the memory exists but in pieces too small for the request. Start Python with expandable segments:
export PYTORCH_ALLOC_CONF=expandable_segments:True
python train.py
PYTORCH_ALLOC_CONF is the current name. The error message still suggests PYTORCH_CUDA_ALLOC_CONF, which PyTorch
keeps as an alias, and the docs mark the option as experimental.
Fixes in order of cost#
Each step down this list costs more than the one before it, in speed, quality or money, so stop at the first one that works.
1. Lower the batch size and keep the effective batch
Activations scale with batch size, so a smaller per-device batch is the first real fix to try. Gradient accumulation keeps the effective batch, here 1 x 16 = 16, by running more, smaller steps before each optimizer update:
args = SFTConfig( # TrainingArguments takes the same two arguments
per_device_train_batch_size = 1,
gradient_accumulation_steps = 16,
# ...your other settings
)
Unsloth gives the same advice for out-of-memory errors: set the batch size to 1, 2 or 3.
2. Shorten the sequence length
Activations also scale with sequence length, and so does the KV cache during inference. In TRL, max_length caps it
(1,024 tokens by default in TRL 1.14), and in Unsloth it is max_seq_length. Check your data's length distribution
before you cut, because truncation drops the end of every longer example. TRL's padding_free option, which works
with FlashAttention 2 or 3, also saves the memory that padding tokens would take.
3. Use lower precision
Mixed-precision training with AdamW costs 18 bytes per parameter in Hugging Face's accounting: 6 for the weights, 8 for the optimizer states and 4 for the gradients. An 8-bit optimizer from bitsandbytes cuts the optimizer part from 8 bytes to 2. Run in bf16 mixed precision as well, which keeps the activations in 16-bit instead of 32-bit. TRL's docs ask for an Ampere or newer GPU for bf16, and every GPU in the QuantaCloud catalog qualifies, since the oldest, the RTX A6000 and A100, are Ampere.
args = SFTConfig(bf16 = True, optim = "adamw_8bit")
4. Turn on gradient checkpointing
Gradient checkpointing keeps only some activations and recomputes the rest in the backward pass. The ZeRO paper's
example is a 1.5B GPT-2 at sequence length 1K and batch size 32: about 60 GB of activations, and about 8 GB with
checkpointing, for 33% more computation. Transformers puts the slowdown at about 20% and Axolotl at about 30%. TRL's
SFTConfig turns it on by default as of 1.14, while plain TrainingArguments leaves it off:
args = TrainingArguments(gradient_checkpointing = True)
# Unsloth: FastModel.get_peft_model(model, use_gradient_checkpointing = "unsloth", ...)
5. Quantize the base model
Loading a frozen base model in 4-bit is the QLoRA move: about half a byte per parameter instead of 2. With Transformers and bitsandbytes:
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
bnb = BitsAndBytesConfig(
load_in_4bit = True,
bnb_4bit_quant_type = "nf4",
bnb_4bit_compute_dtype = torch.bfloat16,
bnb_4bit_use_double_quant = True,
)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config = bnb, dtype = "auto")
If you were running a full fine-tune, moving to LoRA or QLoRA is the bigger step, because gradients and optimizer states then exist only for a small adapter. The LoRA, QLoRA and full fine-tuning guide has the numbers by model size, and the Unsloth tutorial runs a QLoRA job end to end. For inference, load an FP8 or 4-bit checkpoint of the model instead of BF16.
6. Offload to CPU memory
Offload keeps part of the model or the optimizer in system RAM and moves it to the GPU when needed, which is slow and
needs a lot of host memory. For inference with Transformers, device_map="auto" fills the GPUs first and puts the
remaining layers in CPU RAM. For training, FSDP and DeepSpeed can offload parameters: Hugging Face PEFT's 70B QLoRA
run dropped from 35.6 to 19.6 GB per GPU with offload, and used about 107 GB of host RAM to do it. Check the RAM of
your configuration before you count on offload. In vLLM, --cpu-offload-gb does the same for weights.
7. Move to a GPU with more memory
When the job only fits with batch size 1, checkpointing on and a shortened sequence, and it still runs out, stop tuning and move up a tier. The usable column is our calculation at vLLM's default budget of 92%, which I also use as the working budget for training:
| GPU memory | Usable at 92% (our calculation) | QuantaCloud GPUs |
|---|---|---|
| 48 GB | 41.4 GiB with ECC on | RTX A6000, RTX 6000 Ada, L40, L40S |
| 80 GB | 73.6 GiB on the A100, 73.3 GiB on the H100 PCIe | A100 80GB, H100 PCIe |
| 96 GB | 87.9 GiB | RTX PRO 6000 Blackwell |
| 141 GB | 129.2 GiB | H200 NVL |
Live prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
| H100 PCIe | - | Not listed | No |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 19:50 UTC
The rule I follow: if the weights plus the activations or KV cache of one full-size step do not fit in the usable column, move up a tier instead of fighting flags. Past 141 GB, split the job across 2, 4 or 8 GPUs in one VM.
Moving means a new instance, and stopping the old one deletes its disk. Push checkpoints and adapters to the Hugging Face Hub or object storage, or copy them off with scp, before you stop. The deploy docs cover launching the new one.
vLLM out of memory: the startup errors and their flags#
vLLM claims its memory at startup, so its out-of-memory errors usually arrive before the first request and name the setting to change. Version 0.30.0 has five, in the order they can appear:
| Error begins with | What it means | What to change |
|---|---|---|
| Free memory on device ... is less than desired GPU memory utilization | Something else already holds GPU memory, often another server or a notebook | Stop the other process, or lower --gpu-memory-utilization |
| Failed to load model | The weights alone do not fit on the GPU | An FP8 or 4-bit checkpoint, --tensor-parallel-size across more GPUs, or a bigger GPU |
| No available memory for the cache blocks | The weights and the startup profiling fill the whole budget | A higher --gpu-memory-utilization if nothing else shares the GPU, an FP8 or 4-bit checkpoint, --tensor-parallel-size across more GPUs, or a bigger GPU |
| To serve at least one request with the model's max seq len | The KV cache for one full-length request does not fit | --max-model-len at or below the length the message estimates, or --kv-cache-dtype fp8 |
| CUDA out of memory occurred when warming up sampler | The warm-up batch of --max-num-seqs requests does not fit in what is left | Lower --max-num-seqs or --gpu-memory-utilization |
--gpu-memory-utilization defaults to 0.92 and applies per vLLM instance, so raising it helps only when nothing else
shares the GPU. --max-model-len auto picks the longest context that fits. An FP8 KV cache halves the cache, but
without calibration it uses scales of 1.0, so check the output quality on your own prompts. For the last few
gigabytes, --max-num-seqs lowers how many requests vLLM batches at once, --enforce-eager gives up CUDA graphs and
the memory they take at some cost in speed, and multimodal models can skip inputs you do not use with
--limit-mm-per-prompt. A typical fix for a model that almost fits, kept on 127.0.0.1 because vLLM otherwise listens
on the VM's public IP:
vllm serve Qwen/Qwen3-8B \
--host 127.0.0.1 \
--max-model-len 16384 \
--kv-cache-dtype fp8 \
--max-num-seqs 64
After startup, vLLM logs a line that starts with GPU KV cache size: and gives the maximum concurrency at your
--max-model-len, which tells you how much room the flags bought. When you split a model across GPUs without NVLink,
such as the L40S, vLLM's docs recommend pipeline parallelism over tensor parallelism. The
vLLM Docker guide and the vLLM install guide show these flags in
full commands, and the KV cache guide sizes the cache for any model.
My order of operations: read the message, clear anything else off the GPU, then batch size, sequence length, precision and checkpointing. If the job still does not fit in the usable memory of the card, move up a tier rather than stacking workarounds, and use how much VRAM you need to pick it. The GPU catalog lists every tier with its live price.