GPU guide

Set up Open WebUI with Ollama on a cloud GPU

Set up Open WebUI with Ollama on an NVIDIA cloud GPU two ways: the ready template, or a pinned Docker Compose file on Ubuntu. Then save your chats.

Faiz Ahmed9 min read

The fastest way to set up Open WebUI with Ollama on a cloud GPU is QuantaCloud's Open WebUI + Ollama template: launch it, sign in, pull a model and chat. The second way is a Docker Compose file on the Bare Metal template. It takes a few more commands, and in exchange you pin the versions, set the configuration with environment variables and can rebuild the same setup every time.

Both paths run on an Ubuntu 22.04 VM with the NVIDIA driver and Docker, you connect as ubuntu, and in both cases stopping the instance deletes everything on it. The last steps cover getting your chats off first.

TemplateDocker Compose on Bare Metal
What you doLaunch, sign in, pull a modelSSH in, write two files, start them, open a tunnel
VersionsWhatever the template runs that dayPinned: Open WebUI v0.11.4 and Ollama 0.34.4 here
How you open itOpen Application, behind your QuantaCloud loginAn SSH tunnel to port 3000
Configuration by environment variableNoYes

Pick a GPU for your models#

The GPU has to hold the model and its context. Models up to about 30B parameters at 4-bit fit a 48 GB card with room for context: gpt-oss:20b is a 14 GB download in the Ollama library and gemma4:31b is 20 GB. llama3.3:70b (43 GB) and gpt-oss:120b (65 GB) want 80 GB or more, such as an H100 PCIe, an RTX PRO 6000 or an H200 NVL. The Open WebUI page has the full model table, and how much VRAM you need covers the arithmetic.

This guide uses a single RTX A6000 and gpt-oss:20b. Ollama needs NVIDIA driver 550 or newer, and every GPU QuantaCloud offers has compute capability 8.0 or higher, well above Ollama's floor of 5.0.

The template path#

The template path takes six steps and no files.

Launch Open WebUI + Ollama on an RTX A6000
  1. Launch the template. The button opens the deploy page with an RTX A6000 offer and the template selected. If your balance is short, add credit first (the minimum is $5), and when you come back, check that the template still reads Open WebUI + Ollama before you click Deploy.

  2. Wait for Running. The status moves through Initializing, Provisioning and Connecting, and app templates stay in Connecting until the app answers its health check.

  3. Open the app. On the Deployments page, click Open Application. QuantaCloud checks your login and that the deployment is yours, then shows Open WebUI.

  4. Create the Open WebUI admin account. Open WebUI has its own sign-in on top of the QuantaCloud login: on a fresh Open WebUI, the first account becomes the administrator and sign-up then switches itself off.

  5. Pull a model. Type gpt-oss:20b into the model selector of a new chat and confirm the download. The admin can also pull from Settings > Admin > Connections, with Manage next to the Ollama connection.

  6. Send a chat, then check that the model runs on the GPU. SSH in with the key you chose at deploy time and ask Ollama inside the app container:

    Terminal
    ssh -i ~/.ssh/your_key ubuntu@<instance-ip>
    C=$(sudo docker ps -q --filter publish=8080)
    sudo docker exec "$C" ollama ps
    

    The PROCESSOR column should read 100% GPU. A CPU share means the model and its context do not fit in GPU memory, and replies will be slow.

The app address goes through QuantaCloud's login gateway. If you would rather your browser talk to the VM directly, forward Open WebUI's port over SSH from your laptop and open http://localhost:8080:

Terminal
ssh -N -L 8080:127.0.0.1:8080 -i ~/.ssh/your_key ubuntu@<instance-ip>

The Docker Compose path on Bare Metal#

The Compose path gives you pinned versions and your own settings. Bare Metal is QuantaCloud's name for the plain Ubuntu 22.04 template with the NVIDIA driver and Docker, and it is still a VM.

Launch an RTX A6000 with Bare Metal
  1. SSH in, check that Docker can see the GPU, and check that Compose is installed:

    Terminal
    ssh -i ~/.ssh/your_key ubuntu@<instance-ip>
    nvidia-smi
    sudo docker run --rm --gpus all ubuntu:22.04 nvidia-smi
    sudo docker compose version
    

    The second command should print the same GPU table as the first, from inside a container. If docker compose is missing, install the plugin from Docker's repository with sudo apt-get update && sudo apt-get install -y docker-compose-plugin.

  2. Create a folder and a secret key. Open WebUI signs its logins with WEBUI_SECRET_KEY, and without a fixed key every recreated container logs everyone out.

    Terminal
    mkdir -p ~/open-webui && cd ~/open-webui
    echo "WEBUI_SECRET_KEY=$(openssl rand -hex 32)" > .env
    
  3. Write compose.yaml in the same folder:

    YAML
    services:
      ollama:
        image: ollama/ollama:0.34.4
        container_name: ollama
        restart: unless-stopped
        volumes:
          - ollama:/root/.ollama
        deploy:
          resources:
            reservations:
              devices:
                - driver: nvidia
                  count: all
                  capabilities: [gpu]
    
      open-webui:
        image: ghcr.io/open-webui/open-webui:v0.11.4
        container_name: open-webui
        restart: unless-stopped
        depends_on:
          - ollama
        ports:
          - "127.0.0.1:3000:8080"
        volumes:
          - open-webui:/app/backend/data
        environment:
          - OLLAMA_BASE_URL=http://ollama:11434
          - WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY}
          - ENABLE_OPENAI_API=false
          - ENABLE_VERSION_UPDATE_CHECK=false
    
    volumes:
      ollama:
      open-webui:
    

    Four choices in this file matter. The image tags are pinned, because Open WebUI's :main and :latest follow its main branch rather than its newest release. Open WebUI is published on 127.0.0.1 only: Docker publishes ports to the outside world by default, and its published ports bypass ufw. Ollama is not published at all, because its API has no authentication, and Open WebUI reaches it over the Compose network at http://ollama:11434. The last two variables switch off Open WebUI's default OpenAI connection and its version check against GitHub.

  4. Start both containers and follow the log:

    Terminal
    sudo docker compose up -d
    sudo docker compose logs -f open-webui
    

    Press Ctrl+C once the log shows Application startup complete, then check curl -s http://127.0.0.1:3000/health, which should answer {"status":true}.

  5. Open a tunnel from your laptop, then create the admin account:

    Terminal
    ssh -N -L 3000:127.0.0.1:3000 -i ~/.ssh/your_key ubuntu@<instance-ip>
    

    Open http://localhost:3000 and create the admin account straight away. The first account on a fresh Open WebUI becomes the administrator, and sign-up switches itself off once it exists.

  6. Pull a model, send a chat, then check where it runs:

    Terminal
    sudo docker exec ollama ollama pull gpt-oss:20b
    sudo docker exec ollama ollama ps
    

    The model selector works here too, as on the template. ollama ps lists a model only once it has answered, so chat first and look for 100% GPU.

For more on SSH keys and a reusable host entry, see connecting to a GPU server over SSH and the SSH docs.

Settings worth changing on the first day#

Context length is the setting to check first. Open WebUI has a num_ctx parameter in Chat Controls > Advanced Parameters and in each model's settings. It is unset by default, which lets Ollama's own default apply, but switching it on pre-fills 2048 tokens, small enough to cut long chats and tool calls short. Leave it unset, or type a real value such as 16384. Set one when a large model would otherwise spill out of GPU memory: on an 80 GB card, Ollama gives llama3.3:70b its full 131,072-token context, a 42.9 GB cache next to 43 GB of weights (our calculation: 131,072 x 327,680 bytes), and 65536 or less keeps it on the GPU.

On the Compose setup, you can set Ollama's server defaults on the ollama service instead: OLLAMA_CONTEXT_LENGTH for the default context, and OLLAMA_KEEP_ALIVE for how long a model stays loaded after a reply, which is 5 minutes by default. OLLAMA_KEEP_ALIVE=1h spares you a model reload after a pause. Add them under an environment: key on the service and run sudo docker compose up -d again to apply them.

To let other people in, turn on New Sign Ups under Settings > Admin > Authentication. New accounts then wait as pending until you approve them in the Admin Panel. On the template, the app address still opens only for your QuantaCloud account.

Save your chats before you stop#

Stopping deletes the instance and its disk: chats, uploaded files, Open WebUI accounts and pulled models all go. QuantaCloud has no volumes or snapshots, so copy what you want to keep first.

  1. Export the chats. In Open WebUI, open Settings > Data Controls > Export Chats. You get one JSON file with every conversation, and Import Chats in the same place restores them on the next instance.

  2. Copy the data folder if you also want uploads, settings and accounts. Open WebUI keeps them in /app/backend/data inside its container. On the template, the container is the one that publishes port 8080:

    Terminal
    C=$(sudo docker ps -q --filter publish=8080)
    sudo docker cp "$C":/app/backend/data ./open-webui-data
    sudo tar czf open-webui-data.tgz open-webui-data
    

    On the Compose setup, the container has a fixed name:

    Terminal
    sudo docker cp open-webui:/app/backend/data ./open-webui-data
    sudo tar czf open-webui-data.tgz open-webui-data
    
  3. Download the archive from your laptop:

    Terminal
    scp -i ~/.ssh/your_key ubuntu@<instance-ip>:open-webui-data.tgz .
    

Models are not worth backing up: ollama pull fetches them again on the next instance, so a list of the tags you use is enough. Billing stops when you stop, and the unused seconds of the current hour come back to your balance, as the pricing page explains.

Troubleshooting#

The five failures below come down to memory, context or a missed step.

SymptomLikely causeFix
The model list is emptyNo model pulled yet, or Open WebUI cannot reach OllamaPull a model first. On Compose, check sudo docker compose ps and that OLLAMA_BASE_URL is http://ollama:11434
Replies are slowollama ps shows a CPU share, so model and context do not fitUse a smaller model, a shorter context or a larger GPU
Replies stop early or come back blanknum_ctx was switched on at 2048Unset it, or set a real value
Everyone is logged out after an updateWEBUI_SECRET_KEY changedKeep the .env file next to compose.yaml
The deploy page shows Bare Metal after you added creditThe template choice was lost on the way back from checkoutSelect Open WebUI + Ollama again before you click Deploy

My rule: use the template when you want to chat today, and the Compose file when you need pinned versions, settings in environment variables or a setup you rebuild often. Either way, export your chats before you press Stop. To call the same models from your own code, continue with using the Ollama API on a remote GPU. To chat with your own documents, see RAG in Open WebUI.

Keep building

Choose your next step.