Wan 2.2 comes in two sizes for ComfyUI, and they need different GPUs. TI2V-5B is one 10.0 GB model that turns text or an image into a 5-second 720P clip at 24 fps, and it runs on any QuantaCloud GPU. The A14B models, one for text-to-video and one for image-to-video, each pair two 14B experts that ComfyUI's template loads as 14.29 GB fp8 files, and Wan's own code peaks at 59.8 GB for a 720P clip. I would run the 5B on an RTX A6000 and take A14B at 720P to the 96 GB RTX PRO 6000.
The Wan 2.2 models and the files ComfyUI loads#
The files decide most of the setup, because each A14B workflow needs both of its experts and a different VAE from the 5B. These are the files ComfyUI's own Wan 2.2 templates load, from Comfy-Org's repackaged repository:
| Model | What it makes | Parameters | Files the template loads | Total (our calculation) |
|---|---|---|---|---|
| TI2V-5B | Text or image to video, 1280 x 704 at 24 fps | 5B, dense | wan2.2_ti2v_5B_fp16 10.00 GB, umT5-XXL fp8 text encoder 6.74 GB, Wan 2.2 VAE 1.41 GB | 18.14 GB |
| T2V-A14B | Text to video, 480P or 720P at 16 fps | 27B in two 14B experts, 14B active per step | High-noise and low-noise fp8_scaled experts, 14.29 GB each, the same text encoder, Wan 2.1 VAE 0.25 GB | 35.58 GB |
| I2V-A14B | Image to video, 480P or 720P at 16 fps | Same as T2V-A14B | Its own pair of fp8_scaled experts, 14.29 GB each, the same text encoder and VAE | 35.58 GB |
All three are Apache-2.0, which allows commercial use. Wan's model cards add that Wan claims no rights over the videos you generate, and that your use must not involve sharing content that breaks the law, harms people or targets vulnerable groups.
Two extras change the numbers. The A14B templates can switch on lightx2v's 4-step Lightning LoRAs, also Apache-2.0, at 1.23 GB per expert, which cut sampling from 20 steps to 4 and bring the text-to-video set to 38.03 GB. Comfy-Org also publishes the experts at fp16, 28.58 GB each, and ComfyUI's Wan examples rank fp16 above fp8_scaled for quality: loading those makes the text-to-video set 64.14 GB instead of 35.58 GB (our calculation).
How much VRAM Wan 2.2 needs, with and without offloading#
The honest answer comes in two sets of numbers: what Wan measured with its reference code, and what ComfyUI needs, which is far less because ComfyUI streams weights between system RAM and the GPU.
| Wan's own test, one GPU | Resolution | Time per clip | Peak memory |
|---|---|---|---|
| TI2V-5B on an RTX 4090 | 720P | 534.7 s | 22.9 GB |
| T2V-A14B on an A100 or A800 | 480P | 785.7 s | 41.3 GB |
| T2V-A14B on an A100 or A800 | 720P | 2,735.7 s | 59.8 GB |
| T2V-A14B on an H100 or H800 | 720P | 1,041.5 s | 59.8 GB |
Those runs already offload: Wan's single-GPU commands use --offload_model True and --convert_model_dtype, plus --t5_cpu for the 5B. Wan's model cards say the 5B command needs a GPU with at least 24 GB, and the A14B command at 720P one with at least 80 GB.
ComfyUI gets by with much less. Its docs say the 5B "should fit well on 8GB vram with the ComfyUI native offloading", and the note inside its 14B text-to-video template records a 24 GB RTX 4090D running a 640 x 640 clip with the fp8_scaled experts at 84% of memory, in about 536 s for the first run, or in about 108 s with the 4-step LoRA. Offloading has a price, though. Weights that do not fit wait in system RAM and move to the GPU when they are needed, so every step slows down, and the offer's RAM becomes part of the budget: RTX A6000 1x offers had 24 to 64 GB of RAM on 2026-09-27, against 144 GB on the RTX PRO 6000 1x.
The card's architecture changes speed more than memory here. The fp8_scaled experts carry per-layer scaling data, so ComfyUI keeps their weights in fp8 on every GPU. The Ada, Hopper and Blackwell cards have FP8 compute. The Ampere RTX A6000 and A100 do not, so ComfyUI converts each layer to 16-bit only while it computes with it: memory stays at fp8 size, and speed pays for the conversion.
One caution before you read nvidia-smi: ComfyUI's dynamic VRAM system, in its stable releases for NVIDIA on Linux since March 2026, keeps weights on the GPU whenever there is room. A high reading on a 96 GB card shows ComfyUI using the space, not the minimum the model needs.
Which GPU for which Wan 2.2 job#
The rule I follow is to compare Wan's published peak with the card's memory and only lean on offloading when I have to.
| Job | GPU I would use | Why |
|---|---|---|
| TI2V-5B at 720P | RTX A6000, 48 GB | 18.14 GB of files and a 22.9 GB peak in Wan's own test |
| A14B at 480P or the 640 x 640 default | RTX A6000, or an Ada card (RTX 6000 Ada, L40, L40S) for FP8 compute | 35.58 GB of files and a 41.3 GB peak at 480P in Wan's own test |
| A14B at 720P | RTX PRO 6000 Blackwell, 96 GB, or an 80 GB A100 or H100 PCIe | Wan's own code peaked at 59.8 GB, more than a 48 GB card holds without offloading |
| A14B at 720P with the fp16 experts | RTX PRO 6000 Blackwell or H200 NVL, 141 GB | 64.14 GB of files before activations |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
The RTX PRO 6000 Blackwell page and the RTX A6000 page list every configuration with its RAM and disk, and ComfyUI GPU requirements covers the image models too.
Launch ComfyUI and download the Wan 2.2 files#
The quickest route is QuantaCloud's ComfyUI template: launch it on the GPU from the table, wait for Running, and connect over SSH as ubuntu. Running ComfyUI on a cloud GPU walks through the template and the install-it-yourself route.
On the template, ComfyUI runs in a Docker container. These lines find the container by its published port and ask the running ComfyUI where its models folder is. The /internal route belongs to ComfyUI's own frontend and can change between releases, so check the path it prints:
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"
Save the file list as models.txt, keeping the lines for the models you want. None of these repositories is gated, so no token is needed:
# Shared by every Wan 2.2 workflow
text_encoders https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors
# TI2V-5B: 18.14 GB with the text encoder
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors
vae https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan2.2_vae.safetensors
# T2V-A14B: 35.58 GB with the text encoder, both experts required
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors
vae https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors
# optional 4-step Lightning LoRAs for T2V
loras https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors
loras https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_low_noise.safetensors
# I2V-A14B: its own experts, same text encoder and Wan 2.1 VAE
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors
loras https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensors
loras https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensors
Then save the download script as get-models.sh. It skips files that are already there:
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi
while read -r folder url; do
case "$folder" in ""|\#*) continue ;; esac
dest="$MODELS/$folder/$(basename "$url")"
if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
mkdir -p "$MODELS/$folder"
echo "downloading $dest"
curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
mv "$dest.part" "$dest"
done < "${1:-models.txt}"
Run it inside the template's container, or directly on your own install:
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" "$C" bash /tmp/get-models.sh /tmp/models.txt
# on your own install instead:
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt
You will run this on every launch. Stopping a QuantaCloud instance deletes it and its disk, and there are no volumes to keep models between sessions, so keep models.txt with your workflows. Loading models on a new ComfyUI instance covers pinned downloads, gated files and Civitai.
Load the workflow and set resolution and length#
The easiest workflows are the ones ComfyUI ships. Press r so the loaders see the new files, open the Templates icon in the sidebar, search for Wan 2.2, and pick the 5B, 14B text-to-video or 14B image-to-video template. Both 14B templates wrap their loaders and samplers in a subgraph with a switch that turns on the 4-step Lightning LoRAs and changes the sampler to 4 steps and CFG 1.
The table below shows the defaults in ComfyUI's workflow templates on 2026-09-28. The ComfyUI version inside QuantaCloud's template is not published yet, so if a template looks different, read comfyui_version from /system_stats and compare.
| Setting | TI2V-5B template | 14B templates |
|---|---|---|
| Size | 1280 x 704, Wan's 720P size for the 5B | 640 x 640. Wan's own sizes are 1280 x 720 for 720P and 832 x 480 for 480P |
| Length | 121 frames at 24 fps, 5 seconds | 81 frames at 16 fps, 5 seconds |
| Steps | 20 | 20, split 10 on the high-noise expert and 10 on the low-noise one. With Lightning: 4, split 2 and 2 |
| CFG | 5 | 3.5. With Lightning: 1 |
| Sampler and scheduler | uni_pc, simple | euler, simple |
| Shift (ModelSamplingSD3) | 8 | 5 |
Length is counted in frames, not seconds. The 14B templates compute it as seconds times frames per second plus one, which is why 5 seconds is 81 frames at 16 fps, and the 5B's 121 frames are 5 seconds at 24 fps. Keep the frame rate each model was made for. Wan's model cards describe 5-second clips, so a longer one is outside what Wan documents, and memory and time grow with every frame you add.
The 14B templates start at 640 x 640, and ComfyUI's docs say its first-and-last-frame template starts small to spare low-VRAM users. On a 48 GB card or larger, set 1280 x 720 in the subgraph's width and height for 720P. Wan's own code samples 40 steps for A14B and 50 for the 5B, so 20 is ComfyUI's choice, not Wan's. For image-to-video, load your picture in the Load Image node of the 14B image-to-video template, or enable the Load Image node in the 5B template with Ctrl+B. Queue with Ctrl+Enter. The Save Video node writes the clips under output/video in the ComfyUI folder.
What a clip costs#
The arithmetic is short: cost per clip is the seconds per clip times the GPU's hourly price, divided by 3,600. Billing runs from launch to stop, so downloads, model loading and idle time count too. QuantaCloud charges the first hour at launch and each further hour when the previous one is used up, and refunds the unused seconds of the current hour when you stop.
As an example with Wan's own numbers, a 720P T2V-A14B clip took 2,735.7 s on one A100 in Wan's test, which at the A100 SXM4 1x price of $1.50 an hour on 2026-09-27 comes to $1.14 (our calculation: 2,735.7 / 3,600 x $1.50). The A100 starts at $1.49/GPU-hr today. ComfyUI's timings differ from Wan's reference code, so treat that as a rough guide and time your own settings.
Save your clips before you stop#
Copy the videos off the instance before you press Stop, every time, because Stop deletes the disk. On the template, copy them out of the container first, in the same session where you set C and MODELS:
COMFY=$(dirname "$MODELS")
sudo docker cp "$C:$COMFY/output" ~/comfy-output
sudo chown -R ubuntu: ~/comfy-output
Then pull the folder to your own machine with rsync -avP ubuntu@<instance-ip>:comfy-output/ ./comfy-output/, and moving files to and from a GPU server covers the other ways. To queue many clips without the browser, send the same workflow from a script with the ComfyUI API from Python.
Wan 2.2 questions#
Can Wan 2.2 run on a 48 GB GPU?
Yes for the 5B at 720P and for A14B at 480P, going by Wan's own peaks of 22.9 GB and 41.3 GB. A14B at 720P peaked at 59.8 GB in Wan's code, so on a 48 GB card it relies on ComfyUI's offloading, which is slower.
Is Wan 2.2 free for commercial use?
Yes. The weights are Apache-2.0, and Wan claims no rights over the content you generate. Wan's model cards still hold you accountable for your use, which must not involve sharing content that breaks the law, harms people or targets vulnerable groups.
Are Wan 2.5, 2.6, 2.7 and 3.0 open weights?
No. Wan-AI had published no weights for any of them on Hugging Face as of 2026-09-28, and in ComfyUI they run only as paid API nodes. The newest open Wan weights are Wan 2.2 derivatives such as Animate-2 14B.
What happens to my models and videos when I stop?
They are deleted with the instance's disk. Download models again from models.txt on the next launch, and copy clips off first.
My decision rule for Wan 2.2: prototype with the 5B on an RTX A6000, switch to A14B when the 5B's clips fall short on your prompts, and run A14B at 720P on the RTX PRO 6000 Blackwell, where both experts fit with room to spare. Try the Lightning LoRA on your own prompts before you rely on it, and copy clips off before you stop. Open-source video generation models compares Wan with HunyuanVideo, LTX and the rest, and the ComfyUI page lists every GPU with the template.
Launch ComfyUI on an RTX PRO 6000 Blackwell