The simplest way to run a private ChatGPT-style assistant is Open WebUI in front of Ollama, on a GPU VM that only your account can open. QuantaCloud's Open WebUI + Ollama template starts both on an NVIDIA GPU instance: Open WebUI is the chat interface in your browser, Ollama downloads and runs the model, and your prompts, the answers and the files you upload are processed and stored on that VM.
It is not ChatGPT, and it does not try to be. It is a self-hosted chat interface for open-weight models that you pull yourself, such as OpenAI's gpt-oss, Meta's Llama 3.3, Qwen and Gemma, and its answers are as good as the model you choose and the GPU memory you give it. What you get in return is a private AI chat in the plain sense of the word: no model provider sees your conversations, and you decide which model answers.
A single RTX A6000 with 48 GB runs from $0.48/GPU-hr, charged by the hour, with the unused seconds refunded when you stop. Prices checked 6 Oct 2026, 01:40 UTC
What the template gives you#
The template is a chat stack without a model: it starts empty, and the first thing you do is pull one.
| Open WebUI + Ollama template | |
|---|---|
| Chat interface | Open WebUI, opened from the Deployments page at a private address under app.quantacloud.net |
| Model runner | Ollama, running next to Open WebUI on the same VM |
| Models included | None. You pull them from the Ollama library inside Open WebUI |
| Who can open it | Only the QuantaCloud account that launched it, after QuantaCloud login |
| Shell access | SSH as ubuntu to the same VM |
| GPUs | Any QuantaCloud GPU, 48 to 141 GB each. The template does not change the price |
| Time to Running | Most single-GPU VMs are running in about 3 minutes (median), and app templates take longer while the app's container image is pulled |
| Disk | The instance disk, for example 256 GB on a single RTX A6000. It is deleted when you stop |
The templates docs list what every QuantaCloud template contains.
How you sign in#
The app address is locked to your QuantaCloud account. Open Application on the Deployments page shows Open WebUI only after QuantaCloud checks that you are logged in and that the deployment is yours, so nobody else can reach it through that address. It also means a teammate cannot: QuantaCloud accounts are single-user today. Your browser reaches the VM through QuantaCloud's login gateway, and if you want a direct line instead, the setup guide shows how to open Open WebUI over an SSH tunnel.
Open WebUI has its own sign-in too. On a fresh install, its first screen offers Create Admin Account, the first account becomes the administrator, and sign-up then switches itself off until the admin turns it back on. Create that account on first open and store the password somewhere safe.
The best GPU for Ollama holds model and context#
The best GPU for Ollama is the smallest one that holds the model and its context together. Ollama keeps the conversation in a KV cache next to the model weights, and it picks the default context length from the total GPU memory: 4k tokens below about 24 GiB, 32k below about 48 GiB and 256k above that, capped at the model's own maximum. When weights and cache do not fit together, Ollama puts part of the model on the CPU, which its own docs tell you to avoid for speed.
| Model (Ollama tag) | Parameters, quantisation | Download | Smallest GPU I would use | QuantaCloud GPUs that fit |
|---|---|---|---|---|
| gpt-oss:20b | 20.9B, MXFP4 | 14 GB | 48 GB | RTX A6000, RTX 6000 Ada, L40, L40S |
| qwen3.6:27b | 27.3B, Q4_K_M | 18 GB | 48 GB | RTX A6000, RTX 6000 Ada, L40, L40S |
| gemma4:31b | 31.3B, Q4_K_M | 20 GB | 48 GB | RTX A6000, RTX 6000 Ada, L40, L40S |
| llama3.3:70b | 70.6B, Q4_K_M | 43 GB | 80 GB | A100 80GB, H100 PCIe, RTX PRO 6000, H200 NVL |
| gpt-oss:120b | 117B, MXFP4 | 65 GB | 80 GB | A100 80GB, H100 PCIe, RTX PRO 6000, H200 NVL |
| qwen3:235b | 235B, Q4_K_M | 142 GB | Two GPUs | 2 x RTX PRO 6000 (192 GB) or 2 x H200 NVL (282 GB) |
Sizes are from the Ollama library on 2026-09-28. The llama3.3:70b row is the one to read twice. Its weights take about 43 GB, and every token of context adds 320 KiB of KV cache at 16-bit precision, so a 32k context adds about 10.7 GB (our calculation: 32,768 x 327,680 bytes). Together that is more than a 48 GB card holds, so on an RTX A6000 it needs a short context to stay on the GPU. On an 80 GB card, Ollama's automatic context for it is the model's full 131,072 tokens, a 42.9 GB cache on its own (our calculation: 131,072 x 327,680 bytes), so set num_ctx to 65536 or less there, as the setup guide shows. gpt-oss:120b is the opposite case: its cache is small, about 4.8 GB at its full 131k context (our calculation from its attention layout), so the 65 GB download fits an 80 GB card with room left. The KV cache guide has the formula, and the gpt-oss guide covers both gpt-oss sizes in detail. Ollama GPU requirements covers more models on 48, 96 and 141 GB.
Where your chats live, and how to keep them#
Everything lives on the instance disk, and Stop deletes it. There is no pause on QuantaCloud: Stop terminates the instance, and your chats, uploaded files, Open WebUI accounts and every model you pulled go with the disk. There are no volumes or snapshots to restore from.
Before you stop, export your chats. In Open WebUI, open Settings > Data Controls > Export Chats: it downloads every conversation as one JSON file, and Import Chats in the same place brings them back on your next instance. Models need no backup, since pulling gpt-oss:20b again is one step. To keep uploaded documents, settings and accounts as well, copy Open WebUI's data folder off the VM over SSH. The setup guide has the commands.
What leaves the VM is a short list, and you control it. Pulling a model downloads it from the Ollama library. Ollama's cloud models, the tags that end in cloud such as gpt-oss:120b-cloud, run on Ollama's servers rather than on your GPU, so a prompt sent to one of them leaves the VM. Features you switch on, such as web search, call their own providers. A stock Open WebUI also makes three calls of its own: a version check to GitHub, a model-list request to its default OpenAI connection, and an update check for its embedding model on Hugging Face. None of them is a chat, and deleting the OpenAI connection under Settings > Admin > Connections removes the second one.
What it costs#
You pay for the GPU instance by the hour, and the template adds nothing to the price. The first hour is charged when you launch, each further hour when the previous one is used up, and when you stop, the unused seconds of the current hour are refunded.
Two worked examples at the prices on 2026-09-27 (our calculation): two hours of gpt-oss:20b on a single RTX A6000 cost 2 x $0.48 = $0.96, and two hours of gpt-oss:120b on an H200 NVL cost 2 x $3.43 = $6.86. Today's prices:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Keep an eye on the balance. If it cannot cover the next hour, QuantaCloud terminates the instance and deletes its disk, with no chance to export first. A low-balance email goes out when the balance drops below $2, the minimum deposit is $5, and auto top-up is optional. The pricing page has the full rules.
Questions before you launch#
Is this the same as ChatGPT?
No. It is a self-hosted alternative: a chat interface running open-weight models on your GPU VM. It does not run OpenAI's hosted models, and the quality of the answers depends on the model you pull. Open WebUI covers the everyday pieces: a chat history, a model picker, and questions about files you drop into a chat.
Can my team use it?
Not through the app address, which opens only for the QuantaCloud account that launched the instance. For a team, run Open WebUI yourself with Docker Compose on the Bare Metal template, put it behind your own HTTPS proxy, and let Open WebUI's accounts decide who gets in: once you turn on New Sign Ups, new accounts wait as pending until an admin approves them. Read Open WebUI's hardening guide before you open it to anyone.
Can I call the model from my own code?
Yes, through Open WebUI. On this template Ollama's port stays inside the app container, so a script calls Open WebUI's API with an Open WebUI API key, over an SSH tunnel to port 8080. The Ollama API guide shows that route, and a plain Ollama server on the Bare Metal template that you reach on port 11434 without exposing it.
Launch it in five steps#
Five steps take you from the Deploy button to a first answer.
- Pick a GPU from the table above and launch it with the Open WebUI + Ollama template. If you add credit on the way, check that the template still reads Open WebUI + Ollama before you click Deploy. The deploy docs walk through the page.
- Wait for Running on the Deployments page.
- Click Open Application, then create the Open WebUI admin account.
- Type
gpt-oss:20binto the model selector of a new chat and confirm the download. - Chat, and export your chats before you stop.
My rule: if the model you want is 30B-class or smaller, start on a 48 GB RTX A6000, and if it is llama3.3:70b or gpt-oss:120b, go straight to 80 GB or more. Launch the template, pull one model, and export your chats before you press Stop. For the full walk-through, including a pinned Docker Compose setup you control, follow the Open WebUI with Ollama setup guide.
Launch Open WebUI + Ollama