The FLUX model I would start with in ComfyUI is FLUX.1 [schnell]: it is Apache-2.0, so you can use it commercially, and it makes an image in 4 steps. Every QuantaCloud GPU has at least 48 GB, so you can skip the fp8 files built for small cards and run FLUX.1 at full 16-bit precision, which is 34.2 GB of files (our calculation). The GPU only becomes a real decision with FLUX.2 [dev], whose ComfyUI files add up to 71.4 GB.
Which FLUX model fits which GPU#
The rule I follow is to add up every file a FLUX workflow loads: the diffusion model, the text encoders and the VAE. If the total fits in the card's memory, ComfyUI can keep everything on the GPU. If it does not, ComfyUI offloads weights to system RAM and each image takes longer. ComfyUI GPU requirements covers other image and video models.
| Model | License | Files ComfyUI loads | Total (our calculation) | Fits in GPU memory at once |
|---|---|---|---|---|
| FLUX.1 [schnell] | Apache-2.0 | flux1-schnell 23.78 GB, T5-XXL fp16 9.79 GB, CLIP-L 0.25 GB, VAE 0.34 GB | 34.15 GB | Every QuantaCloud GPU |
| FLUX.1 [dev] | FLUX.1 [dev] Non-Commercial License | flux1-dev 23.80 GB plus the same text encoders and VAE | 34.17 GB | Every QuantaCloud GPU |
| FLUX.1 Kontext [dev] (image editing) | FLUX.1 [dev] Non-Commercial License | fp8 model 11.90 GB plus the same text encoders and VAE | 22.27 GB | Every QuantaCloud GPU |
| FLUX.2 [klein] 4B | Apache-2.0 | fp8 model 4.07 GB, Qwen3 4B text encoder 8.04 GB, VAE 0.34 GB | 12.45 GB | Every QuantaCloud GPU |
| FLUX.2 [dev] | FLUX Non-Commercial License | fp8 model 35.46 GB, Mistral text encoder 35.58 GB, VAE 0.34 GB | 71.38 GB | The 96 GB and 141 GB cards, and the 80 GB cards on paper |
The totals are weights only, and activations at 1024 by 1024 come on top.
The card's architecture matters as much as its memory, because ComfyUI only computes in FP8 on GPUs with compute capability 8.9 or higher. On the RTX A6000 and A100, which are Ampere cards without FP8, ComfyUI converts the diffusion model in a plain fp8 file, such as the fp8 FLUX.1 checkpoints, to 16-bit as it loads whenever the 16-bit copy fits. Only the T5 text encoder packed into those checkpoints stays in fp8, so on these cards the file mostly saves download time and disk. Files published as fp8_scaled or fp8mixed, such as the Kontext and FLUX.2 [dev] files in this guide, are different: they carry per-layer scaling data, ComfyUI keeps those layers in fp8 on every GPU, and on Ampere it converts each layer to 16-bit only while computing with it. On the Ada, Hopper and Blackwell cards fp8 weights stay in fp8 and take half the memory. NVFP4 files, such as Black Forest Labs' 9.19 GB NVFP4 build of FLUX.1 [dev], only compute in NVFP4 on Blackwell, which on QuantaCloud means the RTX PRO 6000, and ComfyUI's own blog says they need a cu130 PyTorch build to avoid running up to 2x slower than fp8. Check that pytorch_version in /system_stats ends in +cu130 before you pick one.
| QuantaCloud GPU | Memory | FP8 compute in ComfyUI | What I would run on it |
|---|---|---|---|
| RTX A6000 | 48 GB | No (Ampere) | FLUX.1 at 16-bit, and FLUX.2 [klein] 4B |
| RTX 6000 Ada, L40, L40S | 48 GB | Yes (Ada) | FLUX.1 at 16-bit, or fp8 to leave room for LoRAs and larger batches |
| A100 80GB | 80 GB | No (Ampere) | FLUX.1 at 16-bit. FLUX.2 [dev] fits on paper, as on the H100, because its fp8mixed layers stay in fp8 here too |
| H100 PCIe | 80 GB | Yes (Hopper) | FLUX.2 [dev], which fits on paper with about 13 GiB to spare (our calculation) |
| RTX PRO 6000 Blackwell | 96 GB | Yes, plus NVFP4 | FLUX.2 [dev] as ComfyUI's docs ship it, with room to spare |
| H200 NVL | 141 GB | Yes (Hopper) | FLUX.2 [dev] with room for large batches |
On a 48 GB card, FLUX.2 [dev] only runs with offloading, even with the 18.03 GB fp8 text encoder that brings its total to 53.83 GB (our calculation).
Offloaded weights sit in system RAM, so compare the RAM column on the RTX A6000 or RTX 6000 Ada page with the file sizes before you try it.
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Launch ComfyUI and open a shell#
The quickest route is the ComfyUI template on an RTX A6000 for FLUX.1, or on the RTX PRO 6000 Blackwell for FLUX.2 [dev]. Wait for Running, then connect over SSH as ubuntu. Running ComfyUI on a cloud GPU covers the template and the install-it-yourself route step by step.
On the template, ComfyUI runs in a Docker container. These two lines find the container and ask the running ComfyUI where its models folder is. The /internal route belongs to ComfyUI's own frontend and can change between releases, so check the path it prints:
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"
If you installed ComfyUI yourself with comfy-cli, the folder is ~/comfy/ComfyUI/models and you can skip the container steps below.
Put each file in the right folder#
The file names matter as much as the folders, because the Flux.1 workflows in ComfyUI's Templates sidebar expect exactly flux1-schnell.safetensors or flux1-dev.safetensors in diffusion_models, t5xxl_fp16.safetensors and clip_l.safetensors in text_encoders, and ae.safetensors in vae. Save this list as models.txt, keeping the lines for the models you want:
# FLUX.1 [schnell], Apache-2.0: 34.15 GB
diffusion_models https://huggingface.co/Comfy-Org/flux1-schnell/resolve/main/flux1-schnell.safetensors
text_encoders https://huggingface.co/comfyanonymous/flux_text_encoders/resolve/main/t5xxl_fp16.safetensors
text_encoders https://huggingface.co/comfyanonymous/flux_text_encoders/resolve/main/clip_l.safetensors
vae https://huggingface.co/Comfy-Org/Lumina_Image_2.0_Repackaged/resolve/main/split_files/vae/ae.safetensors
# FLUX.1 [dev], non-commercial: accept the license on its Hugging Face page, then set HF_TOKEN
diffusion_models https://huggingface.co/black-forest-labs/FLUX.1-dev/resolve/main/flux1-dev.safetensors
# FLUX.2 [dev], non-commercial: 71.38 GB, for the 96 GB and 141 GB cards
diffusion_models https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/diffusion_models/flux2_dev_fp8mixed.safetensors
text_encoders https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/text_encoders/mistral_3_small_flux2_bf16.safetensors
# on 48 GB and 80 GB cards, use the 18.03 GB fp8 text encoder instead of the line above:
# text_encoders https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/text_encoders/mistral_3_small_flux2_fp8.safetensors
vae https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors
The VAE line points at the copy of ae.safetensors that ComfyUI's own FLUX examples link to, which needs no token. The FLUX.1 [dev] repository is gated: log in to Hugging Face, accept the license on the model page, create a read token, and export HF_TOKEN=... on the VM before you run the script. Then save the download script as get-models.sh:
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi
while read -r folder url; do
case "$folder" in ""|\#*) continue ;; esac
dest="$MODELS/$folder/$(basename "$url")"
if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
mkdir -p "$MODELS/$folder"
echo "downloading $dest"
curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
mv "$dest.part" "$dest"
done < "${1:-models.txt}"
Run it inside the template's container, or directly on your own install:
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" -e HF_TOKEN="${HF_TOKEN:-}" "$C" bash /tmp/get-models.sh /tmp/models.txt
# on your own install instead:
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt
You will run this on every launch: stopping a QuantaCloud instance deletes it and its disk, and there are no volumes to keep models between sessions. Downloading Hugging Face models fast covers quicker tools for large sets.
Load the workflow and generate#
The easiest workflow is the one ComfyUI ships. Press r so the loaders see the new files, open the Templates icon in the sidebar, search for Flux.1, and pick the Schnell or Dev text-to-image template. If a loader still shows a file you do not have, choose the one you downloaded from its list. ComfyUI's FLUX examples page also has working FLUX workflows embedded in its images: drop one onto the canvas to load it. Installing ComfyUI Manager and custom nodes covers adding nodes on a remote GPU.
| Setting | FLUX.1 [schnell] | FLUX.1 [dev] |
|---|---|---|
| Steps | 4 | 20 |
| Sampler and scheduler | euler, simple | euler, simple |
| CFG | 1.0 | 1.0 |
| Guidance | Not used | 3.5, the default, or set with a FluxGuidance node |
| Size | 1024 x 1024 | 1024 x 1024 |
| Text encoders | DualCLIPLoader, type flux, t5xxl_fp16 and clip_l | Same |
Keep CFG at 1.0, as ComfyUI's FLUX examples say. FLUX.1 [dev] takes its guidance from the conditioning instead (3.5 unless a FluxGuidance node changes it), and any CFG above 1.0 makes ComfyUI run a second, unconditional pass on every step. For FLUX.1 Kontext [dev], put the fp8 Kontext model in diffusion_models, reuse the same text encoders and VAE, and start from the Kontext template, which takes an input image and an edit instruction.
On any of these GPUs you can set weight_dtype in the Load Diffusion Model node to fp8_e4m3fn. That forces fp8 storage even on the Ampere cards, halves the diffusion model's memory, and, in the words of ComfyUI's examples, "might reduce quality a tiny bit".
What the licenses allow#
The honest answer is that the model you pick decides what you may do with it, and the Hugging Face mirror you download from changes nothing. The license travels with the weights.
| Model | License | Use the images commercially | Use the model commercially |
|---|---|---|---|
| FLUX.1 [schnell] | Apache-2.0 | Yes | Yes |
| FLUX.2 [klein] 4B | Apache-2.0 | Yes | Yes |
| FLUX.1 [dev], Kontext [dev], Krea [dev] | FLUX.1 [dev] Non-Commercial License | Yes, but not to train a competing model | Only with a license from Black Forest Labs |
| FLUX.2 [dev], FLUX.2 [klein] 9B | FLUX Non-Commercial License | Yes, but not to train a competing model | Only with a license from Black Forest Labs |
Under the FLUX.1 [dev] license, "commercial" is broad: using the model for revenue-generating activity, in anything that interacts with or affects end users, or to train models for commercial use all need Black Forest Labs' license. Companies may still test and evaluate the model in a non-production environment. The license also asks you to run content filtering or review outputs before you distribute them. Read the license text on the model card before you build anything on a [dev] model, and treat this table as a summary, not legal advice.
My decision rule: FLUX.1 [schnell] for anything commercial or high-volume on a 48 GB card, FLUX.1 [dev] at 16-bit when quality matters and the work is non-commercial or licensed, and FLUX.2 [dev] only on the RTX PRO 6000 Blackwell or larger. Copy your outputs off before you stop the instance, because the disk goes with it. To generate in batches without the browser, queue the same workflow from a script with the ComfyUI API from Python. The ComfyUI page lists every GPU with the template and its live price.