Qwen-Image in ComfyUI is a 20B model whose standard fp8 file is 20.43 GB, and with its 9.38 GB text encoder and VAE the whole set is 30.07 GB, so it fits a 48 GB card at once. I would run it on an RTX 6000 Ada, L40 or L40S, and keep the 40.86 GB bf16 file for the 80 GB and 96 GB cards. Qwen-Image, its 2512 update and the Qwen-Image-Edit models are Apache-2.0. The newer Qwen-Image-2.1 is not: it is research-only.
Which Qwen-Image model to run#
The model I would start with is Qwen-Image-2512, the December 2025 update of the original: Qwen describes it as more realistic, more detailed and better at rendering text than the August release, and it keeps the same size, text encoder, VAE and licence. The family has grown since August 2025:
| Model | Released | What it does | Parameters | Licence |
|---|---|---|---|---|
| Qwen-Image | August 2025 | Text to image, with a focus on rendering text in English and Chinese | 20.4B | Apache-2.0 |
| Qwen-Image-2512 | December 2025 | The same model, updated | 20.4B | Apache-2.0 |
| Qwen-Image-Edit-2511 | December 2025 | Image editing from an instruction, with several input images | 20.4B | Apache-2.0 |
| Qwen-Image-2.1 | September 2026 | A new, smaller generator with native 2048 x 2048 output | 7.1B | Qwen Research License, non-commercial |
Earlier edit releases, Qwen-Image-Edit and Edit-2509, and Qwen-Image-Layered are Apache-2.0 too. Qwen-Image-2.1 changes the terms: its licence allows research and evaluation only and asks for a separate licence from Qwen for anything commercial.
The files and where they go#
The files for every Apache model in the table share one text encoder and one VAE, so only the diffusion model changes between them.
| File | Folder under models | Size | What it is |
|---|---|---|---|
qwen_image_fp8_e4m3fn.safetensors | diffusion_models | 20.43 GB | Qwen-Image in plain fp8, the template's default |
qwen_image_fp8mixed.safetensors | diffusion_models | 20.53 GB | The same model in fp8 with per-layer scaling |
qwen_image_bf16.safetensors | diffusion_models | 40.86 GB | The same model at full 16-bit |
qwen_image_2512_fp8_e4m3fn.safetensors | diffusion_models | 20.43 GB | Qwen-Image-2512 in plain fp8 |
qwen_image_edit_2511_fp8mixed.safetensors | diffusion_models | 20.53 GB | Qwen-Image-Edit-2511, the edit template's default |
qwen_2.5_vl_7b_fp8_scaled.safetensors | text_encoders | 9.38 GB | The Qwen2.5-VL 7B text encoder, also published at 16.58 GB in bf16 |
qwen_image_vae.safetensors | vae | 0.25 GB | The VAE |
Qwen-Image-Lightning-8steps-V1.0.safetensors | loras | 1.70 GB | Optional 8-step LoRA from lightx2v, Apache-2.0 |
With the text encoder and VAE, the fp8 set is 30.07 GB and the bf16 set 50.50 GB, or 57.70 GB with the bf16 text encoder (our calculation).
fp8 or bf16#
The honest answer depends on the card, because ComfyUI treats a plain fp8 file differently on GPUs with and without FP8 compute. On the Ada, Hopper and Blackwell cards it keeps the weights in fp8, so Qwen-Image takes 20.43 GB. On the Ampere RTX A6000 and A100, which cannot compute in FP8, ComfyUI converts a plain fp8 file to 16-bit as it loads whenever the 16-bit copy fits within about 88% of the card's memory, less about 1.2 GiB. The 16-bit copy of Qwen-Image is 40.86 GB, under that limit of about 44 GB on a 48 GB card, so an RTX A6000 ends up holding 40.86 GB of weights from a 20.43 GB file (our calculation from ComfyUI's model_management.py).
The fp8mixed file avoids that on Ampere. It carries per-layer scaling data, and ComfyUI keeps those layers in fp8 on every GPU and expands each one to 16-bit only while it computes, so an RTX A6000 holds 20.53 GB and keeps room for the text encoder. Comfy-Org publishes no fp8mixed file for Qwen-Image-2512, so for 2512 on an RTX A6000, or with any plain fp8 file there, set weight_dtype to fp8_e4m3fn in the Load Diffusion Model node, which forces fp8 storage on Ampere too. On an 80 GB or larger card I would skip both and load the bf16 file, the model at its original precision. FP8 vs FP16 vs BF16 explains the formats.
The text encoder is the other half of the memory. The fp8_scaled Qwen2.5-VL file stays at 9.38 GB on every card, and ComfyUI needs it only to encode the prompt, and can move it off the GPU when the diffusion model needs the room.
| GPU | Memory | What I would load |
|---|---|---|
| RTX 6000 Ada, L40, L40S | 48 GB | The fp8 file: 30.07 GB with text encoder and VAE, all of it in memory at once |
| RTX A6000 | 48 GB | The fp8mixed file, which stays at 20.53 GB instead of converting to 40.86 GB. For 2512, the plain fp8 file with weight_dtype set to fp8_e4m3fn |
| A100 80GB, H100 PCIe | 80 GB | The bf16 file: 50.50 GB with text encoder and VAE |
| RTX PRO 6000 Blackwell | 96 GB | bf16 with the bf16 text encoder: 57.70 GB |
| H200 NVL | 141 GB | bf16 with room for batches or the edit model loaded alongside |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | - | Not listed | No |
ComfyUI's own docs ran the fp8 file on a 24 GB RTX 4090D at 86% of its memory, about 94 s for the first image and 71 s for the second, so every QuantaCloud GPU has at least twice the memory of that test card. The RTX 6000 Ada and RTX A6000 pages list their configurations, and ComfyUI GPU requirements compares Qwen-Image with FLUX, SDXL and the video models.
Launch ComfyUI and download the files#
The quickest route is QuantaCloud's ComfyUI template on an RTX 6000 Ada: launch it, wait for Running, and connect over SSH as ubuntu. Running ComfyUI on a cloud GPU covers the template and the install-it-yourself route.
On the template, ComfyUI runs in a Docker container. These lines find the container by its published port and ask the running ComfyUI where its models folder is. The /internal route belongs to ComfyUI's own frontend and can change between releases, so check the path it prints:
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"
Save the file list as models.txt, keeping the lines you need. None of these repositories is gated:
# Shared by Qwen-Image, 2512 and Edit-2511
text_encoders https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/text_encoders/qwen_2.5_vl_7b_fp8_scaled.safetensors
vae https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/vae/qwen_image_vae.safetensors
# Qwen-Image: plain fp8 on Ada, Hopper and Blackwell cards
diffusion_models https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/diffusion_models/qwen_image_fp8_e4m3fn.safetensors
# on the RTX A6000, fp8mixed instead of the line above:
# diffusion_models https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/diffusion_models/qwen_image_fp8mixed.safetensors
# on 80 GB and larger cards, bf16:
# diffusion_models https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/diffusion_models/qwen_image_bf16.safetensors
loras https://huggingface.co/lightx2v/Qwen-Image-Lightning/resolve/main/Qwen-Image-Lightning-8steps-V1.0.safetensors
# Qwen-Image-2512 (on the RTX A6000, set weight_dtype to fp8_e4m3fn in Load Diffusion Model)
diffusion_models https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/resolve/main/split_files/diffusion_models/qwen_image_2512_fp8_e4m3fn.safetensors
loras https://huggingface.co/lightx2v/Qwen-Image-2512-Lightning/resolve/main/Qwen-Image-2512-Lightning-4steps-V1.0-fp32.safetensors
# Qwen-Image-Edit-2511
diffusion_models https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/resolve/main/split_files/diffusion_models/qwen_image_edit_2511_fp8mixed.safetensors
loras https://huggingface.co/lightx2v/Qwen-Image-Edit-2511-Lightning/resolve/main/Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors
Then save the download script as get-models.sh. It skips files that are already there:
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi
while read -r folder url; do
case "$folder" in ""|\#*) continue ;; esac
dest="$MODELS/$folder/$(basename "$url")"
if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
mkdir -p "$MODELS/$folder"
echo "downloading $dest"
curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
mv "$dest.part" "$dest"
done < "${1:-models.txt}"
Run it inside the template's container, or directly on your own install:
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" "$C" bash /tmp/get-models.sh /tmp/models.txt
# on your own install instead:
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt
Stopping a QuantaCloud instance deletes it and its disk, and there are no volumes, so this download happens on every launch. Keep models.txt with your workflows. Loading models on a new ComfyUI instance covers pinned downloads, gated files and Civitai.
Load the workflow and generate#
The easiest workflows are the ones ComfyUI ships. Press r so the loaders see the new files, open the Templates icon in the sidebar and search for Qwen. The Qwen-Image template loads the fp8 model with Load Diffusion Model, the text encoder with Load CLIP set to type qwen_image, and the VAE with Load VAE. If you downloaded the fp8mixed or bf16 file instead, pick it in Load Diffusion Model.
| Setting | Qwen-Image template | With the Lightning LoRA | Qwen-Image-2512 template | Edit-2511 template |
|---|---|---|---|---|
| Size | 1328 x 1328 | 1328 x 1328 | 1328 x 1328 | Scaled from the input image |
| Steps | 20 | 8 | 50, or 4 with its own 4-step LoRA | 40, or 4 with its 4-step LoRA |
| CFG | 4 | 1 | 4, or 1 with the LoRA | 4, or 1 with the LoRA |
| Sampler and scheduler | euler, simple | euler, simple | euler, simple | euler, simple |
| Shift (ModelSamplingAuraFlow) | 3.1 | 3.1 | 3.1 | 3.1 |
Each template has a switch for its Lightning LoRA inside its subgraph, and flipping it changes the steps and CFG with it. Qwen's own reference code runs Qwen-Image at 50 steps and a true CFG of 4.0, and the example code on its model card uses these sizes for each aspect ratio: 1328 x 1328, 1664 x 928 and 928 x 1664, 1472 x 1140 and 1140 x 1472, 1584 x 1056 and 1056 x 1584. Queue with Ctrl+Enter. Images land in ComfyUI's output folder inside the container.
The settings table above shows the defaults in ComfyUI's workflow templates on 2026-09-28. The ComfyUI version inside QuantaCloud's template is not published yet, so if a template looks different, read comfyui_version from /system_stats and compare.
For editing, the Edit-2511 template takes your images in Load Image nodes and an instruction in the positive prompt, such as the template's own example of changing one material into the texture of a second image.
Save your images before you stop#
Copy your images off the instance before you press Stop, because Stop deletes the disk with every model and output on it. On the template, copy them out of the container first, in the same session where you set C and MODELS:
COMFY=$(dirname "$MODELS")
sudo docker cp "$C:$COMFY/output" ~/comfy-output
sudo chown -R ubuntu: ~/comfy-output
Then pull the folder to your own machine with rsync -avP ubuntu@<instance-ip>:comfy-output/ ./comfy-output/. Moving files to and from a GPU server covers the other ways.
What the licences allow#
The licence travels with the weights, whichever mirror you download them from. Qwen-Image, Qwen-Image-2512 and every Qwen-Image-Edit release are Apache-2.0, so you may use the model and its images commercially, keep the licence and notices when you redistribute the weights, and fine-tune it. The Lightning LoRAs from lightx2v are Apache-2.0 as well.
Qwen-Image-2.1 is the exception. Its Qwen Research License, dated September 20, 2026, defines non-commercial use as research or evaluation only, and commercial use needs a separate licence from Qwen. ComfyUI ships templates for it too, with a 7.26 GB int8 model and a 9.35 GB int8 text encoder, 17.28 GB with the VAE (our calculation), plus a 9.47 GB prompt-rewriting model in the default text-to-image template. Check which model a template loads before you use its output in a product. Treat this section as a summary, not legal advice.
Qwen-Image questions#
How much VRAM does Qwen-Image need?
The fp8 files come to 30.07 GB with the text encoder and VAE, and ComfyUI's own test ran them on a 24 GB card at 86% of memory, which means part of them stayed off the GPU. The bf16 set is 50.50 GB, which fits the 80 GB and larger cards at once.
Can I run Qwen-Image on an RTX A6000?
Yes. Load the fp8mixed file so the weights stay at 20.53 GB. The plain fp8 file also works, but ComfyUI converts it to 16-bit on Ampere cards, which takes 40.86 GB of the card's 48 GB, unless you set weight_dtype to fp8_e4m3fn. That setting is the way to run Qwen-Image-2512 there, because it has no fp8mixed file.
Is Qwen-Image free for commercial use?
Yes for Qwen-Image, Qwen-Image-2512 and Qwen-Image-Edit, which are Apache-2.0. No for Qwen-Image-2.1, which is licensed for research and evaluation only.
Which text encoder does Qwen-Image use?
Qwen2.5-VL 7B, loaded with Load CLIP set to type qwen_image. The fp8_scaled file is 9.38 GB and the bf16 file 16.58 GB, and both Qwen-Image and Qwen-Image-Edit use it.
My decision rule for Qwen-Image: the fp8 file on a 48 GB Ada card for everyday work, the fp8mixed file if you are on an RTX A6000, and bf16 on the 80 GB and larger cards when you want the model at full precision. Copy your images off before you stop, because the disk goes with the instance. For FLUX in the same setup, see FLUX in ComfyUI, and the ComfyUI page lists every GPU with the template and its live price.
Launch ComfyUI on an RTX 6000 Ada