The most reliable way to call the ComfyUI API from Python is the pattern ComfyUI's own examples recommend: export the workflow in API format, POST it to /prompt with a client_id, listen on the /ws websocket until that prompt finishes, then read /history and download each file from /view. On a QuantaCloud instance, point the script at an SSH tunnel to port 8188. The Open Application link is built for a browser: it sits behind your QuantaCloud login and an owner check, which a script has no clean way to pass.
You need ComfyUI running on a GPU instance, a workflow that already works in the browser, and Python 3.10 or newer on your own machine. The ComfyUI template and your own install both work, and running ComfyUI on a cloud GPU covers each.
Launch ComfyUI on an RTX A6000Open an SSH tunnel to port 8188#
The one command that makes the rest work is an SSH tunnel from your machine to the instance. Leave it running in its own terminal:
ssh -N -o ServerAliveInterval=30 -o ExitOnForwardFailure=yes -L 8188:127.0.0.1:8188 ubuntu@<instance-ip>
-L forwards port 8188 on your machine to 127.0.0.1:8188 on the VM, and -N opens no shell. ServerAliveInterval=30 makes ssh send a keepalive whenever the server has been silent for 30 seconds, and ExitOnForwardFailure makes ssh quit instead of carrying on without the forward when your local port 8188 is already taken, for example by a local ComfyUI. On the ComfyUI template, ComfyUI answers at 127.0.0.1:8188 on the VM itself, which is where the app URL's tunnel connects. On your own install, ComfyUI binds 127.0.0.1:8188 by default. The SSH docs show where to find the IP and which key to use.
Check the tunnel from a second terminal:
curl -s http://127.0.0.1:8188/system_stats | python3 -m json.tool | head -n 12
The response names the ComfyUI, Python and PyTorch versions. Write the ComfyUI version down: ComfyUI's docs make no compatibility promise for the server API across releases, so a script belongs to the version it was written against.
The tunnel is also the safe way in. ComfyUI has no authentication, so never expose port 8188 with --listen on a public IP to save yourself a terminal window. Connecting to a cloud GPU covers SSH config and keys if you do this often.
Export the workflow in API format#
The format is the part to get right: POST /prompt accepts only the API format, not the file you get from Save. Build and test the workflow in the browser, then open ComfyUI's File menu and choose Export (API). The browser downloads workflow_api.json.
| Saved workflow (File, Save) | API format (File, Export (API)) | |
|---|---|---|
| Structure | nodes and links arrays, positions, groups | One entry per node id, with class_type and inputs |
| A link between nodes | A link id | ["node id", output index], for example ["30", 1] |
Accepted by POST /prompt | No | Yes |
| Loads back into the browser | Yes | Yes, without the layout |
This is a complete API-format workflow for FLUX.1 [schnell], taken from ComfyUI's own example with a new prompt. It needs flux1-schnell-fp8.safetensors (17.24 GB, Apache-2.0) in models/checkpoints, which is one line of the models.txt in running ComfyUI on a cloud GPU, and the FLUX guide explains the settings. Save it as flux_schnell_api.json:
{
"6": {"class_type": "CLIPTextEncode", "inputs": {"text": "a red lighthouse on a cliff at dusk, film photo", "clip": ["30", 1]}},
"8": {"class_type": "VAEDecode", "inputs": {"samples": ["31", 0], "vae": ["30", 2]}},
"9": {"class_type": "SaveImage", "inputs": {"filename_prefix": "flux_schnell", "images": ["8", 0]}},
"27": {"class_type": "EmptySD3LatentImage", "inputs": {"width": 1024, "height": 1024, "batch_size": 1}},
"30": {"class_type": "CheckpointLoaderSimple", "inputs": {"ckpt_name": "flux1-schnell-fp8.safetensors"}},
"31": {"class_type": "KSampler", "inputs": {"seed": 173805153958730, "steps": 4, "cfg": 1.0, "sampler_name": "euler", "scheduler": "simple", "denoise": 1.0, "model": ["30", 0], "positive": ["6", 0], "negative": ["33", 0], "latent_image": ["27", 0]}},
"33": {"class_type": "CLIPTextEncode", "inputs": {"text": "", "clip": ["30", 1]}}
}
Node 6 holds the prompt text, node 31 is the sampler with its seed, and node 9 saves the image. The ids come from the file, so your own export will use different ones.
The Python client#
The client below is the whole integration: one file, the standard library and two packages. Install them on your own machine:
python3 -m pip install requests websocket-client
Save this as comfy_client.py:
#!/usr/bin/env python3
"""Queue a ComfyUI API-format workflow, follow it over the websocket and save its outputs."""
import argparse
import json
import random
import sys
import uuid
from pathlib import Path
import requests
import websocket # from the websocket-client package
def list_nodes(workflow):
for node_id, node in workflow.items():
title = node.get("_meta", {}).get("title", "")
print(f"{node_id:>6} {node['class_type']:<28} {title}")
def set_inputs(workflow, prompt_node, text, seed):
if prompt_node and text is not None:
workflow[prompt_node]["inputs"]["text"] = text
for node in workflow.values():
for key in ("seed", "noise_seed"):
if key in node.get("inputs", {}):
node["inputs"][key] = seed
def queue(server, workflow, client_id):
r = requests.post(f"http://{server}/prompt",
json={"prompt": workflow, "client_id": client_id}, timeout=30)
if r.status_code != 200:
sys.exit(f"ComfyUI returned HTTP {r.status_code}: {r.text}")
return r.json()["prompt_id"]
def wait(ws, prompt_id):
while True:
frame = ws.recv()
if isinstance(frame, bytes):
continue # binary frames are live previews
msg = json.loads(frame)
kind, data = msg["type"], msg.get("data", {})
if data.get("prompt_id") != prompt_id:
continue # queue status, or another job
if kind == "executing" and data["node"] is None:
return # the whole prompt has finished
if kind == "executing":
print(f" running node {data['node']}")
elif kind == "progress":
print(f" node {data['node']}: step {data['value']} of {data['max']}")
elif kind == "execution_error":
sys.exit(f"node {data['node_id']} ({data['node_type']}) failed: {data['exception_message']}")
elif kind == "execution_interrupted":
sys.exit("the job was interrupted")
def save_outputs(server, prompt_id, out_dir):
history = requests.get(f"http://{server}/history/{prompt_id}", timeout=30).json()[prompt_id]
saved = []
for node_output in history["outputs"].values():
for items in node_output.values():
if not isinstance(items, list):
continue
for item in items:
if not isinstance(item, dict) or item.get("type") != "output":
continue # previews live in the temp folder
params = {"filename": item["filename"], "subfolder": item["subfolder"], "type": "output"}
r = requests.get(f"http://{server}/view", params=params, timeout=300)
r.raise_for_status()
path = out_dir / item["filename"]
path.write_bytes(r.content)
saved.append(path)
return saved
def main():
p = argparse.ArgumentParser()
p.add_argument("workflow", help="API-format workflow JSON")
p.add_argument("prompts", nargs="*", help="one job per prompt text")
p.add_argument("--server", default="127.0.0.1:8188")
p.add_argument("--prompt-node", help="id of the node whose text input gets each prompt")
p.add_argument("--seed", type=int, help="fixed seed; the default is a new random seed per job")
p.add_argument("--out", default="outputs")
p.add_argument("--list", action="store_true", help="print node ids and types, then exit")
args = p.parse_intermixed_args()
workflow = json.loads(Path(args.workflow).read_text())
if args.list:
list_nodes(workflow)
return
out_dir = Path(args.out)
out_dir.mkdir(parents=True, exist_ok=True)
client_id = str(uuid.uuid4())
ws = websocket.create_connection(f"ws://{args.server}/ws?clientId={client_id}", timeout=600)
try:
for text in args.prompts or [None]:
seed = args.seed if args.seed is not None else random.randint(0, 2**63 - 1)
set_inputs(workflow, args.prompt_node, text, seed)
prompt_id = queue(args.server, workflow, client_id)
print(f"queued {prompt_id} with seed {seed}")
wait(ws, prompt_id)
for path in save_outputs(args.server, prompt_id, out_dir):
print(f" saved {path}")
finally:
ws.close()
if __name__ == "__main__":
main()
List the nodes to find the id of your prompt node, then queue one job per prompt:
python3 comfy_client.py flux_schnell_api.json --list
python3 comfy_client.py flux_schnell_api.json --prompt-node 6 \
"a red lighthouse on a cliff at dusk, film photo" \
"a snowy harbor at night, long exposure"
Three details carry the design. The script opens one websocket with a random clientId and sends the same value as client_id with every POST /prompt, because ComfyUI only sends a job's execution events to the socket whose id queued it. It sets a new random seed on every node with a seed or noise_seed input before each job, because a job whose inputs have not changed is served from ComfyUI's cache: no node runs and no new file is written. And it treats the executing message with node set to null as the end of its own prompt, then reads /history/<prompt_id> and fetches every file of type output through /view.
What the websocket tells you#
The rule I follow is to key every decision on prompt_id, because the socket also carries queue updates that belong to nobody in particular. These are the message types ComfyUI sends while a job runs:
| Type | When ComfyUI sends it | Fields worth reading |
|---|---|---|
status | On connect and whenever the queue changes | queue_remaining |
execution_start | The prompt starts | prompt_id |
execution_cached | Nodes are skipped because cached outputs exist | nodes |
executing | A node starts, and once more with node null when the prompt is done | node, prompt_id |
progress | A sampler step, or any node that reports progress | value, max, node |
progress_state | Progress changes, with the state of every node that has started | prompt_id, nodes |
executed | A node returned output for the UI, such as saved images | node, output |
execution_success | Every node finished | prompt_id |
execution_error | A node raised an error | node_id, node_type, exception_message |
execution_interrupted | Someone called /interrupt | node_id |
Binary frames are live previews: an 8-byte header followed by the image. The script skips them, and it ignores message types it does not know, so newer ComfyUI releases that add types do not break it.
Send an input image#
Workflows that start from an image, such as image-to-image or a FLUX Kontext edit, need the file on the server first. Upload it to /upload/image, then put the name ComfyUI returns into your LoadImage node:
import requests
with open("input.png", "rb") as f:
r = requests.post("http://127.0.0.1:8188/upload/image",
files={"image": ("input.png", f, "image/png")},
data={"overwrite": "true"}, timeout=60)
r.raise_for_status()
workflow["10"]["inputs"]["image"] = r.json()["name"] # "10" is your LoadImage node id
The file lands in ComfyUI's input folder on the instance. Without overwrite, ComfyUI keeps the old file and gives the new one a numbered name, unless the two files are identical.
When a job fails#
HTTP 400 from /prompt means the workflow failed validation before anything ran, and the node_errors field names the node and the input. A typical one is value_not_in_list, which is what a missing model file produces: ckpt_name: 'flux1-schnell-fp8.safetensors' not in [...]. Download the file into the right folder and run the job again.
Connection refused means the tunnel is down or ComfyUI is not running. Restart the ssh command, and on the template run sudo docker ps -a --format '{{.ID}} {{.Status}}' on the VM to see whether the ComfyUI container is up.
A dropped tunnel does not stop the job. The prompt keeps running on the server: reconnect, then call /history/<prompt_id>, which returns an empty object until the prompt finishes and the outputs after that.
A job you want gone is two calls away. /interrupt stops the running prompt, and /queue with clear empties everything that has not started:
curl -X POST http://127.0.0.1:8188/interrupt
curl -X POST http://127.0.0.1:8188/queue -H 'Content-Type: application/json' -d '{"clear": true}'
Get the results off before you stop#
SaveImage writes each image to ComfyUI's output folder on the instance, and the script copies it to --out on your machine. Stopping a QuantaCloud instance deletes it and its disk, output folder included, so let the last job finish and check the saved files before you press Stop. Because the unused seconds of the current hour are refunded when you stop, there is no reason to keep an idle instance running after a batch. The billing docs have the details.
My rule is to build and debug a workflow in the browser, export it once, and let the script do the repetition. From there, feeding it prompts from a file or from another program is a few lines of Python. To write the prompts with a local model on the same GPU, see Ollama inside ComfyUI. If you have not launched an instance yet, start with the ComfyUI template on an RTX A6000, and keep the tunnel command from this page next to your workflow.