The fastest way to set up Open WebUI with Ollama on a cloud GPU is QuantaCloud's Open WebUI + Ollama template: launch it, sign in, pull a model and chat. The second way is a Docker Compose file on the Bare Metal template. It takes a few more commands, and in exchange you pin the versions, set the configuration with environment variables and can rebuild the same setup every time.
Both paths run on an Ubuntu 22.04 VM with the NVIDIA driver and Docker, you connect as ubuntu, and in both cases stopping the instance deletes everything on it. The last steps cover getting your chats off first.
| Template | Docker Compose on Bare Metal | |
|---|---|---|
| What you do | Launch, sign in, pull a model | SSH in, write two files, start them, open a tunnel |
| Versions | Whatever the template runs that day | Pinned: Open WebUI v0.11.4 and Ollama 0.34.4 here |
| How you open it | Open Application, behind your QuantaCloud login | An SSH tunnel to port 3000 |
| Configuration by environment variable | No | Yes |
Pick a GPU for your models#
The GPU has to hold the model and its context. Models up to about 30B parameters at 4-bit fit a 48 GB card with room for context: gpt-oss:20b is a 14 GB download in the Ollama library and gemma4:31b is 20 GB. llama3.3:70b (43 GB) and gpt-oss:120b (65 GB) want 80 GB or more, such as an H100 PCIe, an RTX PRO 6000 or an H200 NVL. The Open WebUI page has the full model table, and how much VRAM you need covers the arithmetic.
This guide uses a single RTX A6000 and gpt-oss:20b. Ollama needs NVIDIA driver 550 or newer, and every GPU QuantaCloud offers has compute capability 8.0 or higher, well above Ollama's floor of 5.0.
The template path#
The template path takes six steps and no files.
Launch Open WebUI + Ollama on an RTX A6000-
Launch the template. The button opens the deploy page with an RTX A6000 offer and the template selected. If your balance is short, add credit first (the minimum is $5), and when you come back, check that the template still reads Open WebUI + Ollama before you click Deploy.
-
Wait for Running. The status moves through Initializing, Provisioning and Connecting, and app templates stay in Connecting until the app answers its health check.
-
Open the app. On the Deployments page, click Open Application. QuantaCloud checks your login and that the deployment is yours, then shows Open WebUI.
-
Create the Open WebUI admin account. Open WebUI has its own sign-in on top of the QuantaCloud login: on a fresh Open WebUI, the first account becomes the administrator and sign-up then switches itself off.
-
Pull a model. Type
gpt-oss:20binto the model selector of a new chat and confirm the download. The admin can also pull from Settings > Admin > Connections, with Manage next to the Ollama connection. -
Send a chat, then check that the model runs on the GPU. SSH in with the key you chose at deploy time and ask Ollama inside the app container:
ssh -i ~/.ssh/your_key ubuntu@<instance-ip> C=$(sudo docker ps -q --filter publish=8080) sudo docker exec "$C" ollama psThe PROCESSOR column should read 100% GPU. A CPU share means the model and its context do not fit in GPU memory, and replies will be slow.
The app address goes through QuantaCloud's login gateway. If you would rather your browser talk to the VM directly, forward Open WebUI's port over SSH from your laptop and open http://localhost:8080:
ssh -N -L 8080:127.0.0.1:8080 -i ~/.ssh/your_key ubuntu@<instance-ip>
The Docker Compose path on Bare Metal#
The Compose path gives you pinned versions and your own settings. Bare Metal is QuantaCloud's name for the plain Ubuntu 22.04 template with the NVIDIA driver and Docker, and it is still a VM.
Launch an RTX A6000 with Bare Metal-
SSH in, check that Docker can see the GPU, and check that Compose is installed:
ssh -i ~/.ssh/your_key ubuntu@<instance-ip> nvidia-smi sudo docker run --rm --gpus all ubuntu:22.04 nvidia-smi sudo docker compose versionThe second command should print the same GPU table as the first, from inside a container. If
docker composeis missing, install the plugin from Docker's repository withsudo apt-get update && sudo apt-get install -y docker-compose-plugin. -
Create a folder and a secret key. Open WebUI signs its logins with
WEBUI_SECRET_KEY, and without a fixed key every recreated container logs everyone out.mkdir -p ~/open-webui && cd ~/open-webui echo "WEBUI_SECRET_KEY=$(openssl rand -hex 32)" > .env -
Write
compose.yamlin the same folder:services: ollama: image: ollama/ollama:0.34.4 container_name: ollama restart: unless-stopped volumes: - ollama:/root/.ollama deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] open-webui: image: ghcr.io/open-webui/open-webui:v0.11.4 container_name: open-webui restart: unless-stopped depends_on: - ollama ports: - "127.0.0.1:3000:8080" volumes: - open-webui:/app/backend/data environment: - OLLAMA_BASE_URL=http://ollama:11434 - WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY} - ENABLE_OPENAI_API=false - ENABLE_VERSION_UPDATE_CHECK=false volumes: ollama: open-webui:Four choices in this file matter. The image tags are pinned, because Open WebUI's
:mainand:latestfollow its main branch rather than its newest release. Open WebUI is published on 127.0.0.1 only: Docker publishes ports to the outside world by default, and its published ports bypass ufw. Ollama is not published at all, because its API has no authentication, and Open WebUI reaches it over the Compose network athttp://ollama:11434. The last two variables switch off Open WebUI's default OpenAI connection and its version check against GitHub. -
Start both containers and follow the log:
sudo docker compose up -d sudo docker compose logs -f open-webuiPress Ctrl+C once the log shows
Application startup complete, then checkcurl -s http://127.0.0.1:3000/health, which should answer{"status":true}. -
Open a tunnel from your laptop, then create the admin account:
ssh -N -L 3000:127.0.0.1:3000 -i ~/.ssh/your_key ubuntu@<instance-ip>Open
http://localhost:3000and create the admin account straight away. The first account on a fresh Open WebUI becomes the administrator, and sign-up switches itself off once it exists. -
Pull a model, send a chat, then check where it runs:
sudo docker exec ollama ollama pull gpt-oss:20b sudo docker exec ollama ollama psThe model selector works here too, as on the template.
ollama pslists a model only once it has answered, so chat first and look for 100% GPU.
For more on SSH keys and a reusable host entry, see connecting to a GPU server over SSH and the SSH docs.
Settings worth changing on the first day#
Context length is the setting to check first. Open WebUI has a num_ctx parameter in Chat Controls > Advanced Parameters and in each model's settings. It is unset by default, which lets Ollama's own default apply, but switching it on pre-fills 2048 tokens, small enough to cut long chats and tool calls short. Leave it unset, or type a real value such as 16384. Set one when a large model would otherwise spill out of GPU memory: on an 80 GB card, Ollama gives llama3.3:70b its full 131,072-token context, a 42.9 GB cache next to 43 GB of weights (our calculation: 131,072 x 327,680 bytes), and 65536 or less keeps it on the GPU.
On the Compose setup, you can set Ollama's server defaults on the ollama service instead: OLLAMA_CONTEXT_LENGTH for the default context, and OLLAMA_KEEP_ALIVE for how long a model stays loaded after a reply, which is 5 minutes by default. OLLAMA_KEEP_ALIVE=1h spares you a model reload after a pause. Add them under an environment: key on the service and run sudo docker compose up -d again to apply them.
To let other people in, turn on New Sign Ups under Settings > Admin > Authentication. New accounts then wait as pending until you approve them in the Admin Panel. On the template, the app address still opens only for your QuantaCloud account.
Save your chats before you stop#
Stopping deletes the instance and its disk: chats, uploaded files, Open WebUI accounts and pulled models all go. QuantaCloud has no volumes or snapshots, so copy what you want to keep first.
-
Export the chats. In Open WebUI, open Settings > Data Controls > Export Chats. You get one JSON file with every conversation, and Import Chats in the same place restores them on the next instance.
-
Copy the data folder if you also want uploads, settings and accounts. Open WebUI keeps them in
/app/backend/datainside its container. On the template, the container is the one that publishes port 8080:C=$(sudo docker ps -q --filter publish=8080) sudo docker cp "$C":/app/backend/data ./open-webui-data sudo tar czf open-webui-data.tgz open-webui-dataOn the Compose setup, the container has a fixed name:
sudo docker cp open-webui:/app/backend/data ./open-webui-data sudo tar czf open-webui-data.tgz open-webui-data -
Download the archive from your laptop:
scp -i ~/.ssh/your_key ubuntu@<instance-ip>:open-webui-data.tgz .
Models are not worth backing up: ollama pull fetches them again on the next instance, so a list of the tags you use is enough. Billing stops when you stop, and the unused seconds of the current hour come back to your balance, as the pricing page explains.
Troubleshooting#
The five failures below come down to memory, context or a missed step.
| Symptom | Likely cause | Fix |
|---|---|---|
| The model list is empty | No model pulled yet, or Open WebUI cannot reach Ollama | Pull a model first. On Compose, check sudo docker compose ps and that OLLAMA_BASE_URL is http://ollama:11434 |
| Replies are slow | ollama ps shows a CPU share, so model and context do not fit | Use a smaller model, a shorter context or a larger GPU |
| Replies stop early or come back blank | num_ctx was switched on at 2048 | Unset it, or set a real value |
| Everyone is logged out after an update | WEBUI_SECRET_KEY changed | Keep the .env file next to compose.yaml |
| The deploy page shows Bare Metal after you added credit | The template choice was lost on the way back from checkout | Select Open WebUI + Ollama again before you click Deploy |
My rule: use the template when you want to chat today, and the Compose file when you need pinned versions, settings in environment variables or a setup you rebuild often. Either way, export your chats before you press Stop. To call the same models from your own code, continue with using the Ollama API on a remote GPU. To chat with your own documents, see RAG in Open WebUI.