GPU guide

How to run ComfyUI on a cloud GPU

Run ComfyUI on a rented NVIDIA GPU: pick a GPU for your models, launch the template or install it yourself, load models, and save outputs before you stop.

Faiz Ahmed11 min read

The fastest way to run ComfyUI on a cloud GPU is QuantaCloud's ComfyUI template on a 48 GB RTX A6000: pick the offer, wait for it to boot, and open ComfyUI in your browser behind your QuantaCloud login. Install ComfyUI yourself on the Bare Metal template when you need a specific release or your own custom nodes. Both paths start from an empty disk, so the steps that deserve the most care are getting models on quickly and getting outputs off before you stop.

You need a QuantaCloud account, at least $5 of credit (the minimum deposit), and the list of models your workflow loads.

Pick a GPU for your models#

The rule I follow is to add up the files a workflow loads and compare the total with the card's memory. ComfyUI can offload weights to system RAM when they do not fit, but everything runs slower when it has to. Every GPU QuantaCloud offers has at least 48 GB, which holds each image model below in the files listed.

ModelLicenseFiles you downloadMemory guidance
SDXL 1.0CreativeML Open RAIL++-M, commercial use allowedOne 6.94 GB checkpointStability AI: works on 8 GB cards
FLUX.1 [schnell]Apache-2.0A 17.24 GB fp8 checkpoint, or 34.15 GB of 16-bit filesEither fits a 48 GB card (our calculation)
FLUX.1 [dev]FLUX.1 [dev] Non-Commercial LicenseA 17.25 GB fp8 checkpoint, or 34.17 GB of 16-bit filesEither fits a 48 GB card (our calculation)
Qwen-ImageApache-2.020.43 GB fp8 model plus a 9.38 GB fp8 text encoderComfyUI measured 86% use of a 24 GB card
Qwen-Image-2.1Qwen Research License, non-commercial7.26 GB int8 model plus a 9.35 GB int8 text encoderNo official figure
Wan 2.2 TI2V-5B (video)Apache-2.010.0 GB fp16 modelWan: 24 GB for 720P. ComfyUI: about 8 GB with offloading
Wan 2.2 A14B (video)Apache-2.0Two 14.29 GB fp8 experts, both requiredWan's own code peaks at 41.3 GB (480P) and 59.8 GB (720P)
HunyuanVideo 1.5 (video)Tencent Hunyuan Community License, not licensed in the EU, UK or South KoreaAbout 8.3 GB fp8 modelTencent: 14 GB minimum with offloading

Two licenses in that table need a closer look. FLUX.1 [dev] outputs may be used commercially, but running the model itself for revenue, or in anything end users interact with, needs a commercial license from Black Forest Labs. Qwen-Image-2.1 is research-only, while the original Qwen-Image is Apache-2.0.

For ComfyUI I split QuantaCloud's GPUs into two tiers. The 48 GB cards (RTX A6000, RTX 6000 Ada, L40 and L40S) cover every image model above and the smaller video models. The 96 GB RTX PRO 6000 Blackwell is for Wan 2.2 A14B at 720P without offloading, and for FLUX.2 [dev], whose files come to 71.4 GB. I would start on the RTX A6000 and move to the RTX PRO 6000 Blackwell only when a workflow runs out of memory. How much VRAM you need covers the arithmetic for other models, and ComfyUI GPU requirements lists VRAM by model.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes

Option A: the ComfyUI template#

The template is the shortest path because QuantaCloud starts ComfyUI for you. It runs in a Docker container on the VM, answers on port 8188 and gets a private app URL. The templates docs list what each template contains.

  1. Launch an RTX A6000 with the ComfyUI template. The button opens the deploy page with the offer and the template selected. If you add credit on the way, check that ComfyUI is still the selected template when you return, because the page can come back set to Bare Metal. The deploy docs walk through the page.
Launch ComfyUI on an RTX A6000
  1. Wait for Running. Most single-GPU VMs are running in about 3 minutes (median). App templates take longer, because the deployment stays in Connecting while the app's container image is pulled and until ComfyUI answers its health check.
  1. Click Open Application on the Deployments page. ComfyUI opens at a private app.quantacloud.net address behind your QuantaCloud login, and only the account that owns the deployment can open it.
  2. Record what is in the image. Versions matter for custom nodes and for the API, so read them from the running server over SSH:
Terminal
ssh ubuntu@<instance-ip>
curl -s http://127.0.0.1:8188/system_stats | python3 -m json.tool | head -n 25

The response lists the ComfyUI, Python and PyTorch versions and the GPU memory ComfyUI sees.

  1. Load your models with the script in the models section below. The missing-models prompt that ComfyUI shows when you open a workflow downloads through your browser to your own computer, which does not help on a server. ComfyUI-Manager may not help either: at its default security level it blocks model installs whenever ComfyUI listens beyond localhost, and a ComfyUI inside a container usually does.
  2. Click the Templates icon in ComfyUI's sidebar, pick a workflow that matches your models, and queue it with Ctrl+Enter. Images land in ComfyUI's output folder inside the container.
  3. Copy your outputs off, then stop the instance. Both steps are covered below.

Option B: install ComfyUI yourself on Ubuntu#

Installing ComfyUI yourself is the path for a specific release, your own custom nodes, or a ComfyUI server that only listens on localhost. The Bare Metal template gives you Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect as ubuntu.

Launch an Ubuntu GPU VM on an RTX A6000
  1. Connect and check the driver.
Terminal
ssh ubuntu@<instance-ip>
nvidia-smi

The header shows the driver version and the newest CUDA version it supports. ComfyUI's README requires a cu130 PyTorch build on RTX 20-series and newer GPUs, and cu130 builds need driver R580 or newer.

comfy-cli reads the same value and installs the newest PyTorch build the driver supports. If the header shows a CUDA version below 13.0, the driver predates R580, comfy-cli falls back to an older build that ComfyUI's current README does not support, and the template is the simpler path on that VM. Check your GPU, driver and CUDA version explains the header.

  1. Install comfy-cli in a virtual environment, then install ComfyUI with it.
Terminal
sudo apt-get update && sudo apt-get install -y python3-venv git
python3 -m venv ~/comfy-venv && source ~/comfy-venv/bin/activate
pip install --upgrade pip comfy-cli
comfy --skip-prompt install --nvidia --version latest

Keep --version latest. Without it, comfy-cli installs the nightly master branch, and ComfyUI's README warns that commits between stable tags can break custom nodes. ComfyUI lands in ~/comfy/ComfyUI with ComfyUI-Manager installed.

  1. Start the server in the background and check that it answers.
Terminal
comfy launch --background
curl -s http://127.0.0.1:8188/system_stats | head -c 300

ComfyUI binds 127.0.0.1:8188 unless you pass --listen or --port after --, and comfy-cli turns the Manager on. comfy stop shuts the server down.

  1. From your own machine, open an SSH tunnel and browse to http://127.0.0.1:8188.
Terminal
ssh -N -L 8188:127.0.0.1:8188 ubuntu@<instance-ip>

The one thing I always check on a remote ComfyUI server is that nothing else can reach port 8188. ComfyUI has no authentication: anyone who reaches the port can queue jobs, read your outputs and, through the Manager, install custom nodes. Keep it on 127.0.0.1 and use the tunnel. Never start it with a bare --listen on a public IP. If you run it in Docker, publish it as -p 127.0.0.1:8188:8188, because a plain -p 8188:8188 listens on every interface and Docker's rules bypass ufw. The tunnel also keeps ComfyUI-Manager useful: it treats a loopback server as local, and at its default security level it blocks model installs, updates and uninstalls on a server that listens on other addresses. Connecting to a cloud GPU covers SSH config and VS Code on the same instance.

If you would rather skip comfy-cli, the manual install from ComfyUI's README is a clone and two pip commands, run here in a venv. The --extra-index-url flag keeps PyPI available for everything except PyTorch. This route leaves the Manager out: install manager_requirements.txt and start with --enable-manager if you want it.

Terminal
sudo apt-get update && sudo apt-get install -y python3-venv git
git clone --depth 1 --branch v0.37.0 https://github.com/Comfy-Org/ComfyUI.git ~/ComfyUI
cd ~/ComfyUI && python3 -m venv .venv && source .venv/bin/activate
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python main.py

Get models onto a fresh instance quickly#

The honest answer is that you download your models on every launch. Stopping a QuantaCloud instance deletes it and its disk, and there are no volumes or snapshots to keep them in, so each instance starts empty. What makes that bearable is a model list and a script. Put one folder and one URL per line in models.txt:

Output
# folder under models/   URL
checkpoints   https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/sd_xl_base_1.0.safetensors
checkpoints   https://huggingface.co/Comfy-Org/flux1-schnell/resolve/main/flux1-schnell-fp8.safetensors

Then save this as get-models.sh. It skips files that are already there and sends your Hugging Face token when HF_TOKEN is set. Gated repositories, such as Black Forest Labs' FLUX.1 [dev], need the token after you accept the license on the model page.

Terminal
#!/usr/bin/env bash
# Usage: MODELS=/path/to/ComfyUI/models bash get-models.sh models.txt
set -euo pipefail
MODELS="${MODELS:-$HOME/comfy/ComfyUI/models}"
AUTH=()
if [ -n "${HF_TOKEN:-}" ]; then AUTH=(-H "Authorization: Bearer $HF_TOKEN"); fi

while read -r folder url; do
  case "$folder" in ""|\#*) continue ;; esac
  dest="$MODELS/$folder/$(basename "$url")"
  if [ -s "$dest" ]; then echo "already there: $dest"; continue; fi
  mkdir -p "$MODELS/$folder"
  echo "downloading $dest"
  curl -fL --retry 3 "${AUTH[@]}" -o "$dest.part" "$url"
  mv "$dest.part" "$dest"
done < "${1:-models.txt}"

On your own install, point it at ComfyUI's models folder:

Terminal
MODELS=~/comfy/ComfyUI/models bash get-models.sh models.txt

On the template, run the same script inside the container. The second line asks the running ComfyUI where its model folders are. That /internal route exists for ComfyUI's own frontend and may change between releases, so check the path it prints.

Terminal
C=$(sudo docker ps -q --filter publish=8188)
MODELS=$(curl -s http://127.0.0.1:8188/internal/folder_paths | python3 -c 'import json, os, sys; print(os.path.dirname(json.load(sys.stdin)["checkpoints"][0]))')
echo "$MODELS"
sudo docker cp get-models.sh "$C":/tmp/ && sudo docker cp models.txt "$C":/tmp/
sudo docker exec -e MODELS="$MODELS" -e HF_TOKEN="${HF_TOKEN:-}" "$C" bash /tmp/get-models.sh /tmp/models.txt

The two files above come to 24.17 GB (our calculation: 6,938,078,334 plus 17,236,328,572 bytes).

Press r in ComfyUI afterwards so the loader nodes list the new files. These are the folders most workflows use:

Folder under modelsWhat goes thereNode that loads it
checkpointsAll-in-one files: model, text encoder and VAE togetherLoad Checkpoint
diffusion_modelsDiffusion model weights on their ownLoad Diffusion Model
text_encodersCLIP, T5 and other text encodersLoad CLIP, DualCLIPLoader
vaeVAE filesLoad VAE
lorasLoRA filesLoad LoRA

For larger model sets, downloading Hugging Face models fast compares faster download tools.

Save your work before you stop#

Copy your outputs off the instance before you press Stop, every time. Stop terminates the VM and deletes its disk, and nothing brings it back. On your own install, pull the folders straight to your machine:

Terminal
rsync -avP ubuntu@<instance-ip>:comfy/ComfyUI/output/ ./comfy-output/
rsync -avP ubuntu@<instance-ip>:comfy/ComfyUI/user/default/workflows/ ./comfy-workflows/

On the template, copy them out of the container first. Run this on the VM, in the same session where you set C and MODELS:

Terminal
COMFY=$(dirname "$MODELS")
sudo docker cp "$C:$COMFY/output" ~/comfy-output
sudo docker cp "$C:$COMFY/user/default/workflows" ~/comfy-workflows
sudo chown -R ubuntu: ~/comfy-output ~/comfy-workflows

Then run rsync from your machine against comfy-output/ and comfy-workflows/. A workflow only exists in that folder after you save it in ComfyUI, so save before you copy.

When something goes wrong#

Re-select the ComfyUI template if the deploy page comes back set to Bare Metal. That happens when you add credit in the middle of deploying, and the fix is to pick ComfyUI again before you click Deploy.

Open Application does not load. The deployment stays in Connecting until ComfyUI answers its health check. Over SSH, sudo docker ps -a --format '{{.ID}} {{.Status}}' shows whether the container is up, and sudo docker logs --tail 50 <container-id> shows why it is not.

A loader node does not list your file. The file is in the wrong folder, or ComfyUI has not rescanned yet, so press r. Queued through the API, the same problem comes back as a value_not_in_list error.

Torch not compiled with CUDA enabled, on your own install. A CPU build of PyTorch got in. Run pip uninstall torch, then the cu130 install command from option B.

CUDA out of memory. Set weight_dtype to fp8_e4m3fn in the Load Diffusion Model node, which stores the diffusion model in fp8 and halves its memory on any of these GPUs, at a small cost in quality. A plain fp8 file, such as the fp8 FLUX.1 checkpoints, is not the same thing on the RTX A6000 or A100: those cards have no FP8 compute, so ComfyUI converts its diffusion model to 16-bit as it loads whenever the 16-bit copy fits. Files published as fp8_scaled or fp8mixed carry per-layer scaling data, and ComfyUI keeps those layers in fp8 on every GPU. The next step up is the 96 GB RTX PRO 6000, and fixing CUDA out of memory covers the other options.

A 403 behind your own reverse proxy. ComfyUI returns 403 when the Host header is a loopback address and the Origin does not match it, which is what a proxy sends when it does not pass the original Host through. Forward the original Host and the WebSocket upgrade headers, and do not turn on --enable-cors-header to get around it, because that replaces the check instead of fixing it.

What it costs#

The billing rule is short: QuantaCloud charges the first hour when the instance launches and each further hour when the previous one is used up, and when you stop, the unused seconds of the current hour are refunded. Boot time counts. A session that runs 50 minutes therefore costs 50/60 of the hourly rate (our calculation), not a full hour. Pricing and the billing docs have the full rules.


My decision rule: use the template when you want ComfyUI in a browser within minutes, and install it yourself when the ComfyUI version or the node set has to be exact. Either way, keep models.txt next to your workflows and copy outputs off before you stop. From here, drive the same server from a script with the ComfyUI API from Python, or set up FLUX in ComfyUI. For a different interface, Stable Diffusion WebUI Forge in Docker covers Forge. The ComfyUI page has the live GPU picker.

Keep building

Choose your next step.