GPU guide

Generate video with Wan 2.2 in ComfyUI on a cloud GPU

Run Wan 2.2 video in ComfyUI: the 5B and A14B files and sizes, VRAM with and without offloading, resolution and length settings, and which GPU to rent.

Faiz Ahmed12 min read

Wan 2.2 comes in two sizes for ComfyUI, and they need different GPUs. TI2V-5B is one 10.0 GB model that turns text or an image into a 5-second 720P clip at 24 fps, and it runs on any QuantaCloud GPU. The A14B models, one for text-to-video and one for image-to-video, each pair two 14B experts that ComfyUI's template loads as 14.29 GB fp8 files, and Wan's own code peaks at 59.8 GB for a 720P clip. I would run the 5B on an RTX A6000 and take A14B at 720P to the 96 GB RTX PRO 6000.

The Wan 2.2 models and the files ComfyUI loads#

The files decide most of the setup, because each A14B workflow needs both of its experts and a different VAE from the 5B. These are the files ComfyUI's own Wan 2.2 templates load, from Comfy-Org's repackaged repository:

ModelWhat it makesParametersFiles the template loadsTotal (our calculation)
TI2V-5BText or image to video, 1280 x 704 at 24 fps5B, densewan2.2_ti2v_5B_fp16 10.00 GB, umT5-XXL fp8 text encoder 6.74 GB, Wan 2.2 VAE 1.41 GB18.14 GB
T2V-A14BText to video, 480P or 720P at 16 fps27B in two 14B experts, 14B active per stepHigh-noise and low-noise fp8_scaled experts, 14.29 GB each, the same text encoder, Wan 2.1 VAE 0.25 GB35.58 GB
I2V-A14BImage to video, 480P or 720P at 16 fpsSame as T2V-A14BIts own pair of fp8_scaled experts, 14.29 GB each, the same text encoder and VAE35.58 GB

All three are Apache-2.0, which allows commercial use. Wan's model cards add that Wan claims no rights over the videos you generate, and that your use must not involve sharing content that breaks the law, harms people or targets vulnerable groups.

Two extras change the numbers. The A14B templates can switch on lightx2v's 4-step Lightning LoRAs, also Apache-2.0, at 1.23 GB per expert, which cut sampling from 20 steps to 4 and bring the text-to-video set to 38.03 GB. Comfy-Org also publishes the experts at fp16, 28.58 GB each, and ComfyUI's Wan examples rank fp16 above fp8_scaled for quality: loading those makes the text-to-video set 64.14 GB instead of 35.58 GB (our calculation).

How much VRAM Wan 2.2 needs, with and without offloading#

The honest answer comes in two sets of numbers: what Wan measured with its reference code, and what ComfyUI needs, which is far less because ComfyUI streams weights between system RAM and the GPU.

Wan's own test, one GPUResolutionTime per clipPeak memory
TI2V-5B on an RTX 4090720P534.7 s22.9 GB
T2V-A14B on an A100 or A800480P785.7 s41.3 GB
T2V-A14B on an A100 or A800720P2,735.7 s59.8 GB
T2V-A14B on an H100 or H800720P1,041.5 s59.8 GB

Those runs already offload: Wan's single-GPU commands use --offload_model True and --convert_model_dtype, plus --t5_cpu for the 5B. Wan's model cards say the 5B command needs a GPU with at least 24 GB, and the A14B command at 720P one with at least 80 GB.

ComfyUI gets by with much less. Its docs say the 5B "should fit well on 8GB vram with the ComfyUI native offloading", and the note inside its 14B text-to-video template records a 24 GB RTX 4090D running a 640 x 640 clip with the fp8_scaled experts at 84% of memory, in about 536 s for the first run, or in about 108 s with the 4-step LoRA. Offloading has a price, though. Weights that do not fit wait in system RAM and move to the GPU when they are needed, so every step slows down, and the offer's RAM becomes part of the budget: RTX A6000 1x offers had 24 to 64 GB of RAM on 2026-09-27, against 144 GB on the RTX PRO 6000 1x.

The card's architecture changes speed more than memory here. The fp8_scaled experts carry per-layer scaling data, so ComfyUI keeps their weights in fp8 on every GPU. The Ada, Hopper and Blackwell cards have FP8 compute. The Ampere RTX A6000 and A100 do not, so ComfyUI converts each layer to 16-bit only while it computes with it: memory stays at fp8 size, and speed pays for the conversion.

One caution before you read nvidia-smi: ComfyUI's dynamic VRAM system, in its stable releases for NVIDIA on Linux since March 2026, keeps weights on the GPU whenever there is room. A high reading on a 96 GB card shows ComfyUI using the space, not the minimum the model needs.

Which GPU for which Wan 2.2 job#

The rule I follow is to compare Wan's published peak with the card's memory and only lean on offloading when I have to.

JobGPU I would useWhy
TI2V-5B at 720PRTX A6000, 48 GB18.14 GB of files and a 22.9 GB peak in Wan's own test
A14B at 480P or the 640 x 640 defaultRTX A6000, or an Ada card (RTX 6000 Ada, L40, L40S) for FP8 compute35.58 GB of files and a 41.3 GB peak at 480P in Wan's own test
A14B at 720PRTX PRO 6000 Blackwell, 96 GB, or an 80 GB A100 or H100 PCIeWan's own code peaked at 59.8 GB, more than a 48 GB card holds without offloading
A14B at 720P with the fp16 expertsRTX PRO 6000 Blackwell or H200 NVL, 141 GB64.14 GB of files before activations
GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

The RTX PRO 6000 Blackwell page and the RTX A6000 page list every configuration with its RAM and disk, and ComfyUI GPU requirements covers the image models too.

Launch ComfyUI and download the Wan 2.2 files#

The quickest route is QuantaCloud's ComfyUI template: launch it on the GPU from the table, wait for Running, and connect over SSH as ubuntu. Running ComfyUI on a cloud GPU walks through the template and the install-it-yourself route.

Launch ComfyUI on an RTX A6000 Launch ComfyUI on an RTX PRO 6000 Blackwell

On the template, ComfyUI runs in a Docker container. These lines find the container by its published port and ask the running ComfyUI where its models folder is. The /internal route belongs to ComfyUI's own frontend and can change between releases, so check the path it prints:

Terminal
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"

Save the file list as models.txt, keeping the lines for the models you want. None of these repositories is gated, so no token is needed:

Output
# Shared by every Wan 2.2 workflow
text_encoders    https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors

# TI2V-5B: 18.14 GB with the text encoder
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors
vae              https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan2.2_vae.safetensors

# T2V-A14B: 35.58 GB with the text encoder, both experts required
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors
vae              https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors
# optional 4-step Lightning LoRAs for T2V
loras            https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors
loras            https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_low_noise.safetensors

# I2V-A14B: its own experts, same text encoder and Wan 2.1 VAE
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors
diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors
loras            https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensors
loras            https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensors

Then save the download script as get-models.sh. It skips files that are already there:

Terminal
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi

while read -r folder url; do
  case "$folder" in ""|\#*) continue ;; esac
  dest="$MODELS/$folder/$(basename "$url")"
  if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
  mkdir -p "$MODELS/$folder"
  echo "downloading $dest"
  curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
  mv "$dest.part" "$dest"
done < "${1:-models.txt}"

Run it inside the template's container, or directly on your own install:

Terminal
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" "$C" bash /tmp/get-models.sh /tmp/models.txt

# on your own install instead:
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt

You will run this on every launch. Stopping a QuantaCloud instance deletes it and its disk, and there are no volumes to keep models between sessions, so keep models.txt with your workflows. Loading models on a new ComfyUI instance covers pinned downloads, gated files and Civitai.

Load the workflow and set resolution and length#

The easiest workflows are the ones ComfyUI ships. Press r so the loaders see the new files, open the Templates icon in the sidebar, search for Wan 2.2, and pick the 5B, 14B text-to-video or 14B image-to-video template. Both 14B templates wrap their loaders and samplers in a subgraph with a switch that turns on the 4-step Lightning LoRAs and changes the sampler to 4 steps and CFG 1.

The table below shows the defaults in ComfyUI's workflow templates on 2026-09-28. The ComfyUI version inside QuantaCloud's template is not published yet, so if a template looks different, read comfyui_version from /system_stats and compare.

SettingTI2V-5B template14B templates
Size1280 x 704, Wan's 720P size for the 5B640 x 640. Wan's own sizes are 1280 x 720 for 720P and 832 x 480 for 480P
Length121 frames at 24 fps, 5 seconds81 frames at 16 fps, 5 seconds
Steps2020, split 10 on the high-noise expert and 10 on the low-noise one. With Lightning: 4, split 2 and 2
CFG53.5. With Lightning: 1
Sampler and scheduleruni_pc, simpleeuler, simple
Shift (ModelSamplingSD3)85

Length is counted in frames, not seconds. The 14B templates compute it as seconds times frames per second plus one, which is why 5 seconds is 81 frames at 16 fps, and the 5B's 121 frames are 5 seconds at 24 fps. Keep the frame rate each model was made for. Wan's model cards describe 5-second clips, so a longer one is outside what Wan documents, and memory and time grow with every frame you add.

The 14B templates start at 640 x 640, and ComfyUI's docs say its first-and-last-frame template starts small to spare low-VRAM users. On a 48 GB card or larger, set 1280 x 720 in the subgraph's width and height for 720P. Wan's own code samples 40 steps for A14B and 50 for the 5B, so 20 is ComfyUI's choice, not Wan's. For image-to-video, load your picture in the Load Image node of the 14B image-to-video template, or enable the Load Image node in the 5B template with Ctrl+B. Queue with Ctrl+Enter. The Save Video node writes the clips under output/video in the ComfyUI folder.

What a clip costs#

The arithmetic is short: cost per clip is the seconds per clip times the GPU's hourly price, divided by 3,600. Billing runs from launch to stop, so downloads, model loading and idle time count too. QuantaCloud charges the first hour at launch and each further hour when the previous one is used up, and refunds the unused seconds of the current hour when you stop.

As an example with Wan's own numbers, a 720P T2V-A14B clip took 2,735.7 s on one A100 in Wan's test, which at the A100 SXM4 1x price of $1.50 an hour on 2026-09-27 comes to $1.14 (our calculation: 2,735.7 / 3,600 x $1.50). The A100 starts at $1.49/GPU-hr today. ComfyUI's timings differ from Wan's reference code, so treat that as a rough guide and time your own settings.

Save your clips before you stop#

Copy the videos off the instance before you press Stop, every time, because Stop deletes the disk. On the template, copy them out of the container first, in the same session where you set C and MODELS:

Terminal
COMFY=$(dirname "$MODELS")
sudo docker cp "$C:$COMFY/output" ~/comfy-output
sudo chown -R ubuntu: ~/comfy-output

Then pull the folder to your own machine with rsync -avP ubuntu@<instance-ip>:comfy-output/ ./comfy-output/, and moving files to and from a GPU server covers the other ways. To queue many clips without the browser, send the same workflow from a script with the ComfyUI API from Python.

Wan 2.2 questions#

Can Wan 2.2 run on a 48 GB GPU?

Yes for the 5B at 720P and for A14B at 480P, going by Wan's own peaks of 22.9 GB and 41.3 GB. A14B at 720P peaked at 59.8 GB in Wan's code, so on a 48 GB card it relies on ComfyUI's offloading, which is slower.

Is Wan 2.2 free for commercial use?

Yes. The weights are Apache-2.0, and Wan claims no rights over the content you generate. Wan's model cards still hold you accountable for your use, which must not involve sharing content that breaks the law, harms people or targets vulnerable groups.

Are Wan 2.5, 2.6, 2.7 and 3.0 open weights?

No. Wan-AI had published no weights for any of them on Hugging Face as of 2026-09-28, and in ComfyUI they run only as paid API nodes. The newest open Wan weights are Wan 2.2 derivatives such as Animate-2 14B.

What happens to my models and videos when I stop?

They are deleted with the instance's disk. Download models again from models.txt on the next launch, and copy clips off first.


My decision rule for Wan 2.2: prototype with the 5B on an RTX A6000, switch to A14B when the 5B's clips fall short on your prompts, and run A14B at 720P on the RTX PRO 6000 Blackwell, where both experts fit with room to spare. Try the Lightning LoRA on your own prompts before you rely on it, and copy clips off before you stop. Open-source video generation models compares Wan with HunyuanVideo, LTX and the rest, and the ComfyUI page lists every GPU with the template.

Launch ComfyUI on an RTX PRO 6000 Blackwell

Keep building

Choose your next step.