GPU guide

ComfyUI and Stable Diffusion GPU requirements: VRAM by model

The GPU for Stable Diffusion and ComfyUI depends on the model: the VRAM SD 1.5, SDXL, FLUX, Qwen-Image and Wan 2.2 need, and what offloading changes.

Faiz Ahmed10 min read

The honest answer is that ComfyUI itself needs very little, and the models you load decide the GPU. SD 1.5 and SDXL run on 8 to 10 GB cards. FLUX.1, Qwen-Image and Wan 2.2 5B are sized for 24 GB cards. Wan 2.2 A14B at 720P and FLUX.2 [dev] are the ones their makers size for 80 GB cards. ComfyUI's offloading lowers every one of those minimums, at the cost of speed and system RAM. Every QuantaCloud GPU has at least 48 GB, so the real choice is which card runs your workflow without offloading.

VRAM by model#

The rule I follow is to add up the files a workflow loads and compare the total with the published figure for the model. The file totals below come from the byte sizes on Hugging Face, and the published figures from each model's own docs or ComfyUI's.

WorkflowFiles ComfyUI loads (our calculation)Published VRAM figureQuantaCloud GPU I would start on
SD 1.5One 4.27 GB checkpointAt least 10 GB for the original Stable Diffusion v1 code (CompVis)RTX A6000
SDXL 1.0One 6.94 GB checkpoint8 GB (Stability AI)RTX A6000
FLUX.1 [schnell], fp8 checkpointOne 17.24 GB checkpointNo official figureRTX A6000
FLUX.1 [dev], 16-bit34.17 GB: model, T5-XXL, CLIP-L and VAEModel file about 23 GB, halved by an fp8 weight type (ComfyUI docs)RTX A6000
Qwen-Image, fp830.07 GB: model, text encoder and VAEComfyUI measured 86% use of a 24 GB RTX 4090DRTX 6000 Ada or L40S, or an RTX A6000 with the fp8mixed file
Wan 2.2 TI2V-5B (video)18.14 GB: fp16 model, fp8 text encoder and VAE24 GB for 720P (Wan), about 8 GB with ComfyUI's offloadingRTX A6000
Wan 2.2 T2V-A14B (video), fp835.58 GB: two experts, text encoder and VAE80 GB for 720P, with peaks of 41.3 GB at 480P and 59.8 GB at 720P in Wan's own codeA100 80GB, H100 or RTX PRO 6000 Blackwell
FLUX.2 [dev]71.38 GB: fp8mixed model, bf16 text encoder and VAEAn H100-equivalent GPU for Black Forest Labs' reference scriptRTX PRO 6000 Blackwell or H200 NVL

The totals are weights only. Activations come on top and grow with resolution, frame count and batch size, which is why Wan's figure climbs from 41.3 GB at 480P to 59.8 GB at 720P while the files stay the same. The totals are also a worst case, because ComfyUI does not need every file on the GPU at once: it loads weights when a step needs them and evicts older ones under memory pressure. The FLUX, Qwen-Image and Wan 2.2 guides list every file and where it goes, and loading models on a new instance has a script that downloads a set in one command.

How offloading changes the minimum#

The one feature that changed the answer is DynamicVRAM, which Comfy added to its stable releases for NVIDIA GPUs in March 2026. When a layer's weights do not fit on the GPU, ComfyUI copies them over for that one layer and carries on, instead of stopping with an out-of-memory error. Comfy's own benchmark ran Wan 2.2's two 14B experts, 56 GB of fp16 weights, at 320 x 320 and 81 frames on an RTX 5060, an 8 GB card, with 32 GB and 64 GB of system RAM. In ComfyUI v0.37.0 it is on by default on NVIDIA GPUs with PyTorch 2.8 or newer, and the startup log says DynamicVRAM support detected and enabled.

The memory flags you find in older guides now behave differently:

FlagWhat it does in ComfyUI v0.37.0
NoneDynamicVRAM: weights load when needed and spill to system RAM under pressure
--highvramKeeps models in GPU memory after use, and turns DynamicVRAM off
--gpu-onlyStores and runs everything, text encoders included, on the GPU, and turns DynamicVRAM off
--novramThe option for when low-VRAM mode is not enough, and it turns DynamicVRAM off
--lowvramDoes nothing while DynamicVRAM is on. Without it, runs the text encoders on the CPU
--reserve-vram 8Leaves 8 GB of GPU memory for other software on the same card, such as Ollama
--disable-dynamic-vramReturns to the older loading, which estimates memory before it loads a model

Offloading moves the requirement rather than removing it. Weights that leave the GPU sit in system RAM or are read back from disk, and every copy crosses PCIe, so each step takes longer. On QuantaCloud the configurations differ more in RAM than in GPU memory: in the 2026-09-27 catalog an RTX A6000 1x came with 24 to 64 GB of RAM, an RTX PRO 6000 1x with 144 GB and an H200 NVL 1x with 180 GB. Check the RAM of an offer before you count on offloading a large model.

Precision is the other lever, and the GPU's architecture decides whether it saves memory. ComfyUI computes in FP8 only on compute capability 8.9 or higher, which on QuantaCloud means the RTX 6000 Ada, L40, L40S, H100, H200 NVL and RTX PRO 6000. On the Ampere cards, the RTX A6000 and A100, ComfyUI converts a plain fp8 file to 16-bit as it loads whenever the 16-bit copy fits. Qwen-Image's 20.43 GB fp8_e4m3fn file therefore becomes about 40.86 GB of weights on an A6000, the size of its bf16 file. Files published as fp8_scaled or fp8mixed carry per-layer scaling data and stay fp8 on every GPU, so the 20.53 GB qwen_image_fp8mixed file is the one to download for an A6000. FP8 vs FP16 vs BF16 explains the formats.

Which QuantaCloud GPU runs each workflow#

The rule I follow is to pick the smallest card that holds the whole workflow without offloading, and to treat offloading as a fallback rather than a plan.

QuantaCloud GPUMemoryFP8 compute in ComfyUIWhat I would run on it without offloading
RTX A600048 GBNo (Ampere)SD 1.5, SDXL, FLUX.1 at 16-bit, Qwen-Image from the fp8mixed file, Wan 2.2 5B
RTX 6000 Ada, L40, L40S48 GBYes (Ada)The same, with plain fp8 files at half their 16-bit memory
A100 80GB80 GBNo (Ampere)Wan 2.2 A14B at 720P, and FLUX.2 [dev] on paper because its fp8mixed layers stay fp8
H100 PCIe80 GBYes (Hopper)Wan 2.2 A14B at 720P, and FLUX.2 [dev] on paper with about 13 GiB to spare
RTX PRO 6000 Blackwell96 GBYes, plus NVFP4FLUX.2 [dev] as ComfyUI ships it, and Wan 2.2 A14B with room to spare
H200 NVL141 GBYes (Hopper)Everything above, with room for larger batches and longer clips

Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

The ComfyUI template runs on any of them, and the price is the same with or without it.

Launch ComfyUI on an RTX A6000

Licenses decide more than VRAM#

The license of the model you pick decides what you may do with it, and the Hugging Face mirror you download from changes nothing.

ModelLicenseCommercial use of the model
SDXL 1.0CreativeML Open RAIL++-MAllowed, and its use restrictions pass on to anyone you share it with
FLUX.1 [schnell]Apache-2.0Allowed
FLUX.1 [dev]FLUX.1 [dev] Non-Commercial LicenseOutputs may be used commercially. The model itself needs a license from Black Forest Labs
FLUX.2 [dev]FLUX Non-Commercial LicenseOnly with a license from Black Forest Labs
Qwen-ImageApache-2.0Allowed. The newer Qwen-Image-2.1 is under a research-only license
Wan 2.2Apache-2.0Allowed

Treat the table as a summary and read the license on the model card before you build a product on a model.

Check what your own workflow uses#

The one thing I always check is the peak during the real job, at the resolution and batch size I will actually run. Over SSH, watch the card while ComfyUI generates:

Terminal
nvidia-smi --query-gpu=memory.used,memory.total --format=csv -lms 500

ComfyUI reports its own view, including the system RAM that offloading relies on:

Terminal
curl -s http://127.0.0.1:8188/system_stats | python3 -m json.tool

The response lists vram_total and vram_free for the GPU, ram_total and ram_free for the VM, and the ComfyUI and PyTorch versions. If a workflow still runs out of memory, fixing CUDA out of memory orders the fixes by what they cost, and how much VRAM you need covers LLMs and training on the same cards.

Frequently asked questions#

Can ComfyUI run on 8 GB of VRAM?

Yes, for SD 1.5 and SDXL, and with offloading for much larger models. Comfy's own benchmark ran Wan 2.2 A14B at 320 x 320 on an 8 GB RTX 5060, at the cost of speed and system RAM.

How do I pick a GPU for Stable Diffusion in the cloud?

Start from the files your workflow loads. The card I would rent first is the 48 GB RTX A6000, which holds SD 1.5, SDXL and FLUX.1 at 16-bit with room to spare. Move to an Ada card such as the L40S when you want fp8 files at half their memory, and to the RTX PRO 6000 Blackwell or H200 NVL for FLUX.2 [dev] and Wan 2.2 A14B at 720P.

Does ComfyUI need an NVIDIA GPU?

No. ComfyUI also runs on AMD, Intel Arc and Apple Silicon GPUs, and on the CPU with --cpu, slowly. Every QuantaCloud GPU is NVIDIA.

How much system RAM does ComfyUI need?

There is no official figure. ComfyUI maps model files into memory, so RAM works as a cache: with enough RAM, offloaded weights come back from RAM, and with too little they are read again from disk, which is slower. Match RAM to the models you expect to offload.


My decision rule: add up the files, pick the smallest card that holds them with room for activations, and let offloading rescue a workflow rather than define it. For most people that means ComfyUI on an RTX A6000, an Ada card such as the L40S when fp8 files matter, and the RTX PRO 6000 Blackwell or H200 NVL for FLUX.2 [dev] and Wan 2.2 A14B at 720P. Running ComfyUI on a cloud GPU covers the setup, and the ComfyUI page has the live GPU picker.

Keep building

Choose your next step.