GPU guide

Use Ollama inside ComfyUI for prompt generation on one GPU

Run Ollama next to ComfyUI on one cloud GPU, add the comfyui-ollama node, split GPU memory between them and build a workflow that writes its own prompts.

Faiz Ahmed10 min read

The simplest setup is Ollama and ComfyUI on the same GPU instance, both listening on 127.0.0.1, with the comfyui-ollama custom node sending your short idea to Ollama and passing the prompt it writes to the image model. A 4B model such as gemma3:4b is a 3.3 GB download, so it fits beside FLUX.1 on a single 48 GB card. When memory is tight, the node's keep_alive setting unloads the LLM after each prompt. Neither server has authentication, so neither port leaves the VM: you reach ComfyUI through an SSH tunnel or the template's private app address, and ComfyUI reaches Ollama on the VM itself.

How the pieces connect#

The one thing that has to work is ComfyUI reaching Ollama at http://127.0.0.1:11434, the node's default URL. The ComfyUI server makes that call itself, including when the node fetches its model list, so the address has to work from wherever ComfyUI runs.

RouteComfyUI runsOllama runsHow ComfyUI reaches Ollama
Bare Metal templateOn the VM, installed with comfy-cliOn the VM, as a system service127.0.0.1:11434 on the VM
ComfyUI templateIn its Docker containerIn a second container that shares the ComfyUI container's network127.0.0.1:11434 inside that shared network

The Bare Metal route is the one I would pick: both are ordinary processes, and you control every version. The template route saves you the ComfyUI install, and a Docker feature makes it work without opening any port.

ComfyUI v0.37.0 also has a built-in Generate Text node that runs Gemma and Qwen text models loaded through ComfyUI's own loaders. Ollama is the better fit when you already use its model library, want a model outside those families, or want the same LLM to serve other tools too.

Budget GPU memory for both#

The rule I follow is to add the image model and the LLM together and keep a margin of the card free for activations and the LLM's context. These are the Ollama models I would consider for prompt writing, with their download sizes from ollama.com:

Ollama modelDownloadInputWhat it adds
qwen3:4b-instruct2.5 GBTextThe smallest here, with no thinking step to wait for
gemma3:4b3.3 GBText and imageAccepts an image, so it can describe one
ministral-3:8b6.0 GBText and imageTwice the parameters of the 4B models
gemma3:12b8.1 GBText and imageThree times the parameters of gemma3:4b
mistral-small3.2:24b15 GBText and imageThe model in the node's own example workflows

The image side depends on the GPU's architecture. On an RTX A6000, which has no FP8 compute, ComfyUI converts the diffusion model in the plain fp8 FLUX.1 [schnell] checkpoint to 16-bit as it loads, so budget for FLUX.1 at 16-bit: up to 34.15 GB with its text encoders, and about 5 GB less in practice, because the checkpoint's T5 text encoder stays fp8 (our calculation: half of the 9.79 GB 16-bit T5). With gemma3:4b that is up to 37.45 GB, which leaves at least 10.55 GB of a 48 GB card (our calculation). On an RTX 6000 Ada or L40S the fp8 weights stay fp8, and the same pair is 20.54 GB of files. mistral-small3.2 with FLUX.1 at 16-bit comes to 49.15 GB on the same budget, which leaves no margin on a 48 GB card, and that is where keep_alive 0 earns its place. ComfyUI GPU requirements sizes the other image models the same way.

Three settings do the budgeting:

SettingWhereWhat it does
keep_alive 0 minutesOllama Connectivity nodeOllama unloads the LLM right after each prompt, at the cost of reloading it next time
OLLAMA_CONTEXT_LENGTH=8192Ollama's environmentCaps the context. Ollama's default grows with GPU memory, to 32k tokens or more on these cards, and prompt writing needs a small fraction of that
--reserve-vramComfyUI's launch flagsLeaves that many GB of GPU memory to other software, such as Ollama

While the model is loaded, ollama ps shows how much memory it takes, how it splits between GPU and CPU, and its context.

Set it up on the Bare Metal template#

This route runs everything as ordinary processes on the VM. Launch an RTX A6000 with the Bare Metal template and connect over SSH as ubuntu.

Launch an RTX A6000 with Bare Metal
  1. Install ComfyUI with comfy-cli. Running ComfyUI on a cloud GPU explains each flag and the driver check.
Terminal
sudo apt-get update && sudo apt-get install -y python3-venv git zstd
python3 -m venv ~/comfy-venv && source ~/comfy-venv/bin/activate
pip install --upgrade pip comfy-cli
comfy --skip-prompt install --nvidia --version latest
  1. Install Ollama, pinned to the current release. The install script needs zstd, installed in the previous step, and it sets Ollama up as a system service on 127.0.0.1:11434.
Terminal
curl -fsSL https://ollama.com/install.sh | OLLAMA_VERSION=0.34.4 sh
  1. Cap Ollama's context. Run sudo systemctl edit ollama, add these two lines, then restart the service:
ini
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Terminal
sudo systemctl daemon-reload && sudo systemctl restart ollama
ollama pull gemma3:4b
  1. Install the custom node at a pinned version, download the image model and start ComfyUI. The FLUX.1 [schnell] checkpoint is Apache-2.0 and needs no token.
Terminal
comfy node install comfyui-ollama@2.1.0
mkdir -p ~/comfy/ComfyUI/models/checkpoints
curl -fL -o ~/comfy/ComfyUI/models/checkpoints/flux1-schnell-fp8.safetensors \
  https://huggingface.co/Comfy-Org/flux1-schnell/resolve/main/flux1-schnell-fp8.safetensors
comfy launch --background

Like every custom node, comfyui-ollama is Python code that runs with ComfyUI's permissions, so read it before you install it: the registry release comes from its author, stavsap, and its requirements.txt asks only for the ollama client and dotenv. Installing ComfyUI-Manager and custom nodes covers version pinning and the Manager's security levels.

If memory gets tight, stop ComfyUI with comfy stop and start it again with comfy launch --background -- --reserve-vram 8, setting the number to the size ollama ps reports for your model, rounded up.

  1. From your own machine, open the tunnel and browse to http://127.0.0.1:8188:
Terminal
ssh -N -L 8188:127.0.0.1:8188 ubuntu@<instance-ip>

Or use the ComfyUI template#

On the ComfyUI template, ComfyUI already runs in a Docker container, and the VM publishes only its port 8188, on 127.0.0.1. An Ollama service on the VM would sit outside that container's network. Docker can start a second container inside the first one's network instead, and then 127.0.0.1:11434 means the same thing to both. Launch the template and wait for Running. You open ComfyUI with Open Application on the Deployments page, behind your QuantaCloud login, and run the Docker commands over SSH.

Launch ComfyUI on an RTX A6000
  1. Install comfyui-ollama from ComfyUI-Manager's Custom Nodes Manager, if the template has the Manager, then click Restart. Without the Manager, use the Bare Metal route above. Installing ComfyUI-Manager and custom nodes explains what the Manager allows on the template.
  1. Once ComfyUI is back, start Ollama on the ComfyUI container's network and pull the model. Docker with NVIDIA GPUs explains the --gpus flag.
Terminal
C=$(sudo docker ps -q --filter publish=8188)
sudo docker run -d --name ollama --gpus all --network container:"$C" \
  -e OLLAMA_HOST=127.0.0.1:11434 -e OLLAMA_CONTEXT_LENGTH=8192 \
  -v ollama:/root/.ollama ollama/ollama:0.34.4
sudo docker exec ollama ollama pull gemma3:4b

OLLAMA_HOST=127.0.0.1:11434 overrides the image's default of 0.0.0.0, so Ollama answers only inside the shared network. No -p flag is needed, and Docker does not allow one in this mode.

The Ollama container and its models go when the instance stops, like everything else on the disk. If the ComfyUI container restarts, remove the Ollama container with sudo docker rm -f ollama and run step 2 again, because it shares that container's network.

Build the prompt workflow#

The workflow is FLUX.1 [schnell] with two nodes in front of the positive prompt. Start from ComfyUI's own FLUX.1 [schnell] fp8 checkpoint example: drop its image from ComfyUI's FLUX examples page onto the canvas, and the workflow loads with its settings of 4 steps and CFG 1. The FLUX guide explains them.

  1. Double-click the canvas, search for Ollama Connectivity and add it. Keep the URL http://127.0.0.1:11434, pick gemma3:4b in the model list, and set keep_alive to 0 with the unit in minutes if memory is tight. If the list is empty, fix the connection and press Reconnect on the node.
  2. Add Ollama Generate and connect the Connectivity node's connection output to its connectivity input. Leave think off and the format on text.
  3. Put your short idea in the prompt field and the instruction in system, for example: "You write prompts for an image model. Rewrite the idea as one detailed prompt of about 60 words covering subject, setting, lighting and style. Reply with the prompt only."
  4. Connect Ollama Generate's result output to the text input of the positive CLIP Text Encode (Prompt) node.
  5. Add a Preview as Text node and connect result to it as well, so you can read what the LLM wrote.
  6. Queue the workflow with Ctrl+Enter.

ComfyUI runs a node again only when its inputs change, so queuing the same idea again reuses the cached prompt instead of asking Ollama. For a new prompt from the same idea, add an Ollama Options node, connect it to the options input of Ollama Generate, switch enable_seed on and change the seed.

To drive the same workflow from a script, export it in API format as the ComfyUI API guide shows and change the Ollama Generate node's prompt input per job. The client in that guide writes to a text input, so point it at prompt for this node.

When something goes wrong#

The model list on Ollama Connectivity is empty. ComfyUI cannot reach Ollama at the URL on the node. On Bare Metal, check systemctl status ollama and curl http://127.0.0.1:11434/api/version. On the template, check that the Ollama container is running with sudo docker ps, then press Reconnect.

The node fails with a model not found error. The model was never pulled on this instance, and every new instance starts empty, so pull it again.

The first prompt after a pause is slow. Ollama unloaded the model, after 5 idle minutes by default or at once with keep_alive 0. A longer keep_alive keeps it in memory, at the cost of the memory it holds.

ollama ps shows part of the model on the CPU. The card was too full when Ollama loaded it. Lower the context, pick a smaller model, or start ComfyUI with --reserve-vram.

Frequently asked questions#

Do I need a second GPU for Ollama?

No. A 4B model and FLUX.1 at 16-bit fit on one 48 GB card together, by our calculation above. A second GPU only pays off when both models are large, and the Ollama GPU requirements cover the bigger LLMs.

Can the LLM write a prompt from an image?

Yes, with a model that accepts image input, such as gemma3:4b. Connect a Load Image node to the images input of Ollama Generate and ask it to describe the picture as a prompt.

Is Ollama's API reachable from the internet in this setup?

No. Ollama listens on 127.0.0.1, on the VM or inside the shared container network, and its API has no authentication. Using the Ollama API on a remote GPU covers safe remote access when you need it.

Do the models stay when I stop the instance?

No. Stopping a QuantaCloud instance deletes its disk, Ollama's models and ComfyUI's included. Keep your pull list and your workflow file, and set them up again on the next launch.


My rule for Ollama in ComfyUI: one GPU, both servers on 127.0.0.1, the smallest LLM that writes prompts you like, and keep_alive 0 whenever the two models together crowd the card. Start on an RTX A6000 with the Bare Metal route, and move to an Ada card such as the L40S when you want FLUX.1's fp8 weights to stay fp8. The ComfyUI page launches the template on every GPU.

Keep building

Choose your next step.