The honest answer is a sum: the model's weights at the precision you run them, plus the working memory the job builds up, plus room for the runtime. Working memory is the KV cache when you serve an LLM, the text encoders and activations when you generate images or video, and the gradients, optimizer states and activations when you train. Add them up for your model and settings, then round up to the next GPU. On QuantaCloud that means 48, 80, 96 or 141 GB per GPU, with 2, 4 or 8 GPUs in one VM when a single card is not enough.
If you are choosing a graphics card for games, none of this applies: the figures are for AI workloads.
The short answer by workload#
The table maps common jobs to the four memory sizes. Fits means the figure is below the card's memory with room left for context or batch. Tight means it fits only with a short context or little room to spare. Reduced means it runs only with lower-precision files, offloading or tiling. No means one card is not enough.
| Workload | Memory it needs | 48 GB | 80 GB | 96 GB | 141 GB |
|---|---|---|---|---|---|
| 8B LLM in BF16, one 32k-token request | about 21 GB (our calculation) | Fits | Fits | Fits | Fits |
| gpt-oss-20b | within 16 GB (OpenAI) | Fits | Fits | Fits | Fits |
| 32B LLM with FP8 weights | 34.3 GB file (Qwen3-32B) plus KV cache | Tight | Fits | Fits | Fits |
| 32B LLM in BF16, one 32k-token request | about 74 GB (our calculation) | No | Tight | Fits | Fits |
| gpt-oss-120b | a single 80 GB GPU (OpenAI) | No | Tight | Fits | Fits |
| 70B LLM with 4-bit weights | 39.8 GB file (Llama 3.3 70B AWQ) plus KV cache | Tight | Fits | Fits | Fits |
| 70B LLM in FP8, full 128k context | about 115.6 GB (our calculation) | No | No | No | Fits |
| 70B LLM in BF16 | 141 GB of weights alone | No | No | No | No |
| SDXL image generation | 8 GB (Stability AI) | Fits | Fits | Fits | Fits |
| FLUX.1 [dev], original 16-bit file | about 23 GB plus text encoders (ComfyUI) | Fits | Fits | Fits | Fits |
| Qwen-Image, FP8 files | 20.4 GB model plus 9.38 GB text encoder | Fits | Fits | Fits | Fits |
| FLUX.2 [dev], FP8 files | 35.46 GB model plus 18.03 GB text encoder | Reduced | Fits | Fits | Fits |
| FLUX.2 [dev], BF16, everything loaded | about 100 GB (our calculation) | No | No | Tight | Fits |
| Wan2.2 TI2V-5B video, 720P | 22.9 GB peak (Wan) | Fits | Fits | Fits | Fits |
| HunyuanVideo, 544x960, 129 frames | 45 GB (Tencent) | Tight | Fits | Fits | Fits |
| Wan2.2 A14B video, 720P | 59.8 GB peak, at least 80 GB advised (Wan) | Reduced | Fits | Fits | Fits |
| HunyuanVideo, 720x1280, 129 frames | 60 GB (Tencent) | Reduced | Fits | Fits | Fits |
| 8B fine-tune, QLoRA | 6 to 14 GB (published tables) | Fits | Fits | Fits | Fits |
| 8B fine-tune, 16-bit LoRA | 16 to 24 GB (published tables) | Fits | Fits | Fits | Fits |
| 32B fine-tune, QLoRA | 24 to 32 GB (published tables) | Fits | Fits | Fits | Fits |
| 32B fine-tune, 16-bit LoRA | 64 to 80 GB (published tables) | No | Tight | Fits | Fits |
| 70B fine-tune, QLoRA | 40 to 48 GB (published tables) | Tight | Fits | Fits | Fits |
| 70B fine-tune, 16-bit LoRA | 160 to 164 GB (published tables) | No | No | No | No |
| 8B full fine-tune with AdamW | 128 to 144 GB before activations (our calculation) | No | No | No | Tight |
The rows marked No on every card run on more than one GPU in the same VM: 2x H200 NVL holds 282 GB, for example. The sections below explain where each figure comes from.
The rule of thumb for LLM inference#
The VRAM an LLM needs for serving is weights plus KV cache plus overhead, and the weights are the easy part: parameters times bytes per parameter. That is 2 bytes in BF16 or FP16, 1 in FP8 and about half a byte in 4-bit, so NVIDIA's own example puts a 7B model at about 14 GB in FP16. Two corrections matter. For mixture-of-experts models, count the total parameters, not the active ones, because every expert has to sit in memory. And real 4-bit files run larger than the arithmetic: gpt-oss-120b is 65.2 GB on disk against 58.4 GB for its parameter count at half a byte. For a dense 70B model such as Llama 3.3 70B, the arithmetic gives 141 GB in BF16, 70.6 GB in FP8 and 35.3 GB in 4-bit (our calculation from its 70.55 billion parameters), and the last two are lower bounds: the published FP8 and 4-bit AWQ files are 72.7 and 39.8 GB, because the embeddings and the output layer stay 16-bit and the quantized layers carry scales.
The KV cache is the part that grows. Per token it is 2 x layers x KV heads x head dimension x bytes per element,
with the numbers taken from the model's config.json. Qwen3-8B stores 144 KiB per token in BF16 (our calculation:
2 x 36 x 8 x 128 x 2 bytes), so one 32,768-token request holds 4.5 GiB, and with its 16.4 GB of BF16 weights that is
the 21 GB in the table. Llama 3.3 70B stores 320 KiB per token, so one full 128k context holds 42.9 GB, and with the
FP8 file the total is about 115.6 GB (72.7 + 42.9 GB), which one H200 NVL holds. Multiply the cache by the number of
requests you serve at once, and halve it with an FP8 KV cache. The KV cache guide works through
more models.
Overhead covers the CUDA context, activations and the serving engine's workspace. EleutherAI's rule of thumb for
inference is about 1.2 times the model's memory. vLLM works differently: by default it claims 92% of the GPU at
startup and fills whatever the weights leave with KV cache, which is why a vLLM server looks almost full in
nvidia-smi from the first second.
Put together for Qwen3-32B in BF16: 65.5 GB of weights plus 8 GiB of KV cache for one 32,768-token request is about 74 GB, or 69.0 GiB, before activations (our calculation). That is tight on an 80 GB card at vLLM's default 92%, with 4.3 to 4.6 GiB to spare, comfortable on the 96 GB RTX PRO 6000, and borderline on a 48 GB card once the weights are FP8, a 34.3 GB file. FP8 math needs compute capability 8.9 or newer, which the RTX 6000 Ada, L40, L40S, H100, H200 NVL and RTX PRO 6000 have and the RTX A6000 and A100 do not, although vLLM still runs FP8 checkpoints on those two as weight-only. The gpt-oss GPU requirements and vLLM with Docker guides apply this to real servers. DeepSeek GPU requirements and Ollama GPU requirements do the same for DeepSeek models and for Ollama.
Image generation#
Image models need the diffusion model, its text encoders and a VAE at the precision you load them, plus activations that grow with resolution and batch size. The text encoders are easy to forget and can be the biggest part: the FLUX.2 [dev] transformer has 32B parameters, and its text encoder, built on a 24B Mistral model, is a 35.6 GB file in BF16. Loaded together in BF16 that is about 100 GB (our calculation: 32 billion x 2 bytes plus 35.6 GB), which only the 141 GB H200 NVL holds with room to spare. The 96 GB RTX PRO 6000 reports 97,887 MiB and would have under 3 GiB left.
Precision is the main lever. FLUX.1 [dev]'s original file is about 23 GB according to ComfyUI's docs, and loading it with an FP8 weight type halves that. ComfyUI can also stream weights between system RAM and the GPU, which runs big models in much less VRAM at the cost of speed and system RAM. My rule for images: add up the files you load at the precision you run, then leave room for the resolution and batch size you want. The ComfyUI guide and the FLUX guide cover the setup, and the ComfyUI template launches it ready to use. ComfyUI GPU requirements lists VRAM by model, and Stable Diffusion WebUI Forge in Docker covers Forge.
Video generation#
Video is where activations take over. The same Wan2.2 A14B model peaks at 41.3 GB at 480P and 59.8 GB at 720P in Wan's own single-GPU tests, and those figures already use its offload flags. Wan advises at least 80 GB for the 720P A14B run. Tencent lists HunyuanVideo at 45 GB for 544x960 and 60 GB for 720x1280, both at 129 frames. The smaller models are the 48 GB options: Wan2.2 TI2V-5B peaks at 22.9 GB at 720P. ComfyUI's docs also describe running HunyuanVideo in about 8 GB with FP8 weights and tiled VAE decoding. Techniques like these, and the offloading described above, are what the Reduced cells in the table rely on. Open-source video generation models lists the models and the GPUs they need.
Fine-tuning#
Training holds far more than the weights: a gradient and an optimizer state for every trainable parameter, plus the activations saved for the backward pass. Full fine-tuning with mixed-precision AdamW costs 16 to 18 bytes per parameter before activations, so even an 8B model needs 128 to 144 GB (our calculation). LoRA freezes the base at 2 bytes per parameter and trains a small adapter, and QLoRA stores the frozen base in 4-bit, about half a byte per parameter. That is why the published tables put a 70B QLoRA run at 40 to 48 GB and a 16-bit LoRA run at 160 to 164 GB.
Sequence length moves all of it. In Unsloth's benchmark, a 70B QLoRA run could train on 12,106-token sequences on 48 GB and 89,389 on 80 GB. The LoRA, QLoRA and full fine-tuning guide has the full tables with their conditions, and the Unsloth tutorial runs a QLoRA job on an RTX A6000 from start to finish.
What each memory tier gets you on QuantaCloud#
Each tier is a different instance, not different code: the same job moves up by launching a bigger GPU.
| GPU memory | GPUs | Jobs it holds on one GPU |
|---|---|---|
| 48 GB | RTX A6000, RTX 6000 Ada, L40, L40S | 8B LLMs in BF16 with long context, 32B with FP8 weights, gpt-oss-20b, FLUX.1 and Qwen-Image, Wan2.2 5B, QLoRA up to 32B |
| 80 GB | A100 80GB, H100 PCIe | gpt-oss-120b, 70B with 4-bit weights and long context, Wan2.2 A14B at 720P, 70B QLoRA |
| 96 GB | RTX PRO 6000 Blackwell | 32B in BF16 with room for context, 32B LoRA, 70B in FP8 at shorter context |
| 141 GB | H200 NVL | 70B in FP8 with a full 128k context, FLUX.2 in BF16 fully loaded, 70B QLoRA with long sequences |
Live prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
| H100 PCIe | - | Not listed | No |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 20:14 UTC
There is no 64 GB tier. A job that needs 50 to 70 GB, like a 32B model's BF16 weights at 65.5 GB, belongs on an 80 or 96 GB card, or split across two 48 GB GPUs. Past 141 GB, split the job across the GPUs of one VM: on 2026-09-27 the largest configurations were 2x H200 NVL (282 GB), 4x RTX PRO 6000 (384 GB), 8x L40S and 8x RTX 6000 Ada (384 GB each) and 8x A100 SXM4 (640 GB). vLLM splits a model with tensor parallelism, and on GPUs without NVLink, such as the L40S, L40, RTX 6000 Ada and RTX PRO 6000, its docs recommend pipeline parallelism instead. For jobs larger than one VM, or for GPUs that are not on demand such as the B200, we build reserved capacity to order: Send a capacity brief.
The live catalog, with counts and regions:
| GPU | Memory | From | GPUs per VM now | Regions | Launch |
|---|---|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | 1x, 2x, 4x, 8x | us-midwest-1, us-midwest-2, us-midwest-3 | Launch |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | 1x, 2x, 4x, 8x | us-midwest-1, us-midwest-2 | Launch |
| L40 | 48 GB | $0.94/GPU-hr | 1x, 2x, 4x | us-midwest-2 | Launch |
| L40S | 48 GB | $1.09/GPU-hr | 1x, 2x, 4x, 8x | us-midwest-1, us-midwest-2 | Launch |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | 2x | us-midwest-2 | Launch |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | 1x, 8x | us-east-1, us-midwest-1, us-midwest-2 | Launch |
Check how much VRAM a job uses#
The one thing I always check is the peak during the real job, not the number at idle. On a Linux GPU server,
nvidia-smi shows it per GPU, refreshed every second:
nvidia-smi --query-gpu=name,memory.used,memory.total --format=csv -l 1
Inside a PyTorch program, ask for the peak directly:
import torch
print(f"{torch.cuda.max_memory_allocated() / 2**30:.1f} GiB peak allocated")
print(f"{torch.cuda.max_memory_reserved() / 2**30:.1f} GiB peak reserved")
The two disagree for a reason: PyTorch keeps memory it has freed in a cache for reuse, so nvidia-smi shows more in
use than your tensors need. If the peak sits at the edge of the card, the
CUDA out of memory guide orders the fixes by what they cost.
My rule: size for the peak of the real job, at the context length or resolution you actually need, and pick the smallest tier that holds it with room to spare. When in doubt, start on a 48 GB card, measure, and move up one tier only when the numbers say so. The GPU catalog lists every tier with its live price.