GPU guide

Run FLUX in ComfyUI on a cloud GPU

Which FLUX model fits a 48, 80 or 96 GB GPU, where each file goes in ComfyUI, the settings that work, and what the FLUX dev and schnell licenses allow.

Faiz Ahmed9 min read

The FLUX model I would start with in ComfyUI is FLUX.1 [schnell]: it is Apache-2.0, so you can use it commercially, and it makes an image in 4 steps. Every QuantaCloud GPU has at least 48 GB, so you can skip the fp8 files built for small cards and run FLUX.1 at full 16-bit precision, which is 34.2 GB of files (our calculation). The GPU only becomes a real decision with FLUX.2 [dev], whose ComfyUI files add up to 71.4 GB.

Which FLUX model fits which GPU#

The rule I follow is to add up every file a FLUX workflow loads: the diffusion model, the text encoders and the VAE. If the total fits in the card's memory, ComfyUI can keep everything on the GPU. If it does not, ComfyUI offloads weights to system RAM and each image takes longer. ComfyUI GPU requirements covers other image and video models.

ModelLicenseFiles ComfyUI loadsTotal (our calculation)Fits in GPU memory at once
FLUX.1 [schnell]Apache-2.0flux1-schnell 23.78 GB, T5-XXL fp16 9.79 GB, CLIP-L 0.25 GB, VAE 0.34 GB34.15 GBEvery QuantaCloud GPU
FLUX.1 [dev]FLUX.1 [dev] Non-Commercial Licenseflux1-dev 23.80 GB plus the same text encoders and VAE34.17 GBEvery QuantaCloud GPU
FLUX.1 Kontext [dev] (image editing)FLUX.1 [dev] Non-Commercial Licensefp8 model 11.90 GB plus the same text encoders and VAE22.27 GBEvery QuantaCloud GPU
FLUX.2 [klein] 4BApache-2.0fp8 model 4.07 GB, Qwen3 4B text encoder 8.04 GB, VAE 0.34 GB12.45 GBEvery QuantaCloud GPU
FLUX.2 [dev]FLUX Non-Commercial Licensefp8 model 35.46 GB, Mistral text encoder 35.58 GB, VAE 0.34 GB71.38 GBThe 96 GB and 141 GB cards, and the 80 GB cards on paper

The totals are weights only, and activations at 1024 by 1024 come on top.

The card's architecture matters as much as its memory, because ComfyUI only computes in FP8 on GPUs with compute capability 8.9 or higher. On the RTX A6000 and A100, which are Ampere cards without FP8, ComfyUI converts the diffusion model in a plain fp8 file, such as the fp8 FLUX.1 checkpoints, to 16-bit as it loads whenever the 16-bit copy fits. Only the T5 text encoder packed into those checkpoints stays in fp8, so on these cards the file mostly saves download time and disk. Files published as fp8_scaled or fp8mixed, such as the Kontext and FLUX.2 [dev] files in this guide, are different: they carry per-layer scaling data, ComfyUI keeps those layers in fp8 on every GPU, and on Ampere it converts each layer to 16-bit only while computing with it. On the Ada, Hopper and Blackwell cards fp8 weights stay in fp8 and take half the memory. NVFP4 files, such as Black Forest Labs' 9.19 GB NVFP4 build of FLUX.1 [dev], only compute in NVFP4 on Blackwell, which on QuantaCloud means the RTX PRO 6000, and ComfyUI's own blog says they need a cu130 PyTorch build to avoid running up to 2x slower than fp8. Check that pytorch_version in /system_stats ends in +cu130 before you pick one.

QuantaCloud GPUMemoryFP8 compute in ComfyUIWhat I would run on it
RTX A600048 GBNo (Ampere)FLUX.1 at 16-bit, and FLUX.2 [klein] 4B
RTX 6000 Ada, L40, L40S48 GBYes (Ada)FLUX.1 at 16-bit, or fp8 to leave room for LoRAs and larger batches
A100 80GB80 GBNo (Ampere)FLUX.1 at 16-bit. FLUX.2 [dev] fits on paper, as on the H100, because its fp8mixed layers stay in fp8 here too
H100 PCIe80 GBYes (Hopper)FLUX.2 [dev], which fits on paper with about 13 GiB to spare (our calculation)
RTX PRO 6000 Blackwell96 GBYes, plus NVFP4FLUX.2 [dev] as ComfyUI's docs ship it, with room to spare
H200 NVL141 GBYes (Hopper)FLUX.2 [dev] with room for large batches

On a 48 GB card, FLUX.2 [dev] only runs with offloading, even with the 18.03 GB fp8 text encoder that brings its total to 53.83 GB (our calculation).

Offloaded weights sit in system RAM, so compare the RAM column on the RTX A6000 or RTX 6000 Ada page with the file sizes before you try it.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

Launch ComfyUI and open a shell#

The quickest route is the ComfyUI template on an RTX A6000 for FLUX.1, or on the RTX PRO 6000 Blackwell for FLUX.2 [dev]. Wait for Running, then connect over SSH as ubuntu. Running ComfyUI on a cloud GPU covers the template and the install-it-yourself route step by step.

Launch ComfyUI on an RTX A6000 Launch ComfyUI on an RTX PRO 6000 Blackwell

On the template, ComfyUI runs in a Docker container. These two lines find the container and ask the running ComfyUI where its models folder is. The /internal route belongs to ComfyUI's own frontend and can change between releases, so check the path it prints:

Terminal
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"

If you installed ComfyUI yourself with comfy-cli, the folder is ~/comfy/ComfyUI/models and you can skip the container steps below.

Put each file in the right folder#

The file names matter as much as the folders, because the Flux.1 workflows in ComfyUI's Templates sidebar expect exactly flux1-schnell.safetensors or flux1-dev.safetensors in diffusion_models, t5xxl_fp16.safetensors and clip_l.safetensors in text_encoders, and ae.safetensors in vae. Save this list as models.txt, keeping the lines for the models you want:

Output
# FLUX.1 [schnell], Apache-2.0: 34.15 GB
diffusion_models https://huggingface.co/Comfy-Org/flux1-schnell/resolve/main/flux1-schnell.safetensors
text_encoders    https://huggingface.co/comfyanonymous/flux_text_encoders/resolve/main/t5xxl_fp16.safetensors
text_encoders    https://huggingface.co/comfyanonymous/flux_text_encoders/resolve/main/clip_l.safetensors
vae              https://huggingface.co/Comfy-Org/Lumina_Image_2.0_Repackaged/resolve/main/split_files/vae/ae.safetensors

# FLUX.1 [dev], non-commercial: accept the license on its Hugging Face page, then set HF_TOKEN
diffusion_models https://huggingface.co/black-forest-labs/FLUX.1-dev/resolve/main/flux1-dev.safetensors

# FLUX.2 [dev], non-commercial: 71.38 GB, for the 96 GB and 141 GB cards
diffusion_models https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/diffusion_models/flux2_dev_fp8mixed.safetensors
text_encoders    https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/text_encoders/mistral_3_small_flux2_bf16.safetensors
# on 48 GB and 80 GB cards, use the 18.03 GB fp8 text encoder instead of the line above:
# text_encoders  https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/text_encoders/mistral_3_small_flux2_fp8.safetensors
vae              https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors

The VAE line points at the copy of ae.safetensors that ComfyUI's own FLUX examples link to, which needs no token. The FLUX.1 [dev] repository is gated: log in to Hugging Face, accept the license on the model page, create a read token, and export HF_TOKEN=... on the VM before you run the script. Then save the download script as get-models.sh:

Terminal
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi

while read -r folder url; do
  case "$folder" in ""|\#*) continue ;; esac
  dest="$MODELS/$folder/$(basename "$url")"
  if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
  mkdir -p "$MODELS/$folder"
  echo "downloading $dest"
  curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
  mv "$dest.part" "$dest"
done < "${1:-models.txt}"

Run it inside the template's container, or directly on your own install:

Terminal
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" -e HF_TOKEN="${HF_TOKEN:-}" "$C" bash /tmp/get-models.sh /tmp/models.txt

# on your own install instead:
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt

You will run this on every launch: stopping a QuantaCloud instance deletes it and its disk, and there are no volumes to keep models between sessions. Downloading Hugging Face models fast covers quicker tools for large sets.

Load the workflow and generate#

The easiest workflow is the one ComfyUI ships. Press r so the loaders see the new files, open the Templates icon in the sidebar, search for Flux.1, and pick the Schnell or Dev text-to-image template. If a loader still shows a file you do not have, choose the one you downloaded from its list. ComfyUI's FLUX examples page also has working FLUX workflows embedded in its images: drop one onto the canvas to load it. Installing ComfyUI Manager and custom nodes covers adding nodes on a remote GPU.

SettingFLUX.1 [schnell]FLUX.1 [dev]
Steps420
Sampler and schedulereuler, simpleeuler, simple
CFG1.01.0
GuidanceNot used3.5, the default, or set with a FluxGuidance node
Size1024 x 10241024 x 1024
Text encodersDualCLIPLoader, type flux, t5xxl_fp16 and clip_lSame

Keep CFG at 1.0, as ComfyUI's FLUX examples say. FLUX.1 [dev] takes its guidance from the conditioning instead (3.5 unless a FluxGuidance node changes it), and any CFG above 1.0 makes ComfyUI run a second, unconditional pass on every step. For FLUX.1 Kontext [dev], put the fp8 Kontext model in diffusion_models, reuse the same text encoders and VAE, and start from the Kontext template, which takes an input image and an edit instruction.

On any of these GPUs you can set weight_dtype in the Load Diffusion Model node to fp8_e4m3fn. That forces fp8 storage even on the Ampere cards, halves the diffusion model's memory, and, in the words of ComfyUI's examples, "might reduce quality a tiny bit".

What the licenses allow#

The honest answer is that the model you pick decides what you may do with it, and the Hugging Face mirror you download from changes nothing. The license travels with the weights.

ModelLicenseUse the images commerciallyUse the model commercially
FLUX.1 [schnell]Apache-2.0YesYes
FLUX.2 [klein] 4BApache-2.0YesYes
FLUX.1 [dev], Kontext [dev], Krea [dev]FLUX.1 [dev] Non-Commercial LicenseYes, but not to train a competing modelOnly with a license from Black Forest Labs
FLUX.2 [dev], FLUX.2 [klein] 9BFLUX Non-Commercial LicenseYes, but not to train a competing modelOnly with a license from Black Forest Labs

Under the FLUX.1 [dev] license, "commercial" is broad: using the model for revenue-generating activity, in anything that interacts with or affects end users, or to train models for commercial use all need Black Forest Labs' license. Companies may still test and evaluate the model in a non-production environment. The license also asks you to run content filtering or review outputs before you distribute them. Read the license text on the model card before you build anything on a [dev] model, and treat this table as a summary, not legal advice.


My decision rule: FLUX.1 [schnell] for anything commercial or high-volume on a 48 GB card, FLUX.1 [dev] at 16-bit when quality matters and the work is non-commercial or licensed, and FLUX.2 [dev] only on the RTX PRO 6000 Blackwell or larger. Copy your outputs off before you stop the instance, because the disk goes with it. To generate in batches without the browser, queue the same workflow from a script with the ComfyUI API from Python. The ComfyUI page lists every GPU with the template and its live price.

Keep building

Choose your next step.