Open WebUI's knowledge bases are the whole RAG setup: you upload files into a collection, Open WebUI splits them into chunks and turns each chunk into a vector with an embedding model, and when you ask a question it adds the closest chunks to the prompt of the model you chat with. On QuantaCloud's Open WebUI + Ollama template all of it runs on one GPU VM, so your documents, their vectors and both models stay on the instance. Two settings decide most of the result: the embedding model, which you choose before the first upload, and the chat model's context length.
The instance disk is also the catch. Stopping the instance deletes it, knowledge base included, so the last section shows how to take the knowledge base with you and load it on the next instance.
What you need#
You need the Open WebUI + Ollama template on a GPU with room for two models, the chat model and a small embedding model. A single RTX A6000 with 48 GB, from $0.48/GPU-hr, holds gpt-oss:20b, a 14 GB download in the Ollama library, and bge-m3, a 1.2 GB embedding model, with most of the card left for context. The setup guide covers launching the template, creating the Open WebUI admin account and pulling a chat model, and this guide starts from there.
Launch Open WebUI + Ollama on an RTX A6000Choose the embedding model before the first upload#
The embedding model is the setting that is expensive to change later. Every chunk is stored as a vector from that model, and Open WebUI's docs say a new embedding model needs a re-index of every knowledge base, because vectors from two models do not compare. Files you dropped straight into a chat are not re-indexed at all: you upload them again.
Out of the box, Open WebUI embeds with sentence-transformers/all-MiniLM-L6-v2 through its built-in SentenceTransformers engine. It is a small model, 22.7 million parameters in a 90.9 MB file, and it reads at most 256 word pieces of each chunk and ignores the rest. It also runs inside the Open WebUI process, not in Ollama, and on the template that process has a CPU-only build of PyTorch. Ollama gets the GPU, and the default embedding model runs on the VM's vCPUs, 6 to 12 of them on the RTX A6000 VMs listed on 2026-09-27.
The fix is to hand embedding to Ollama, which already runs on the GPU. My pick is bge-m3: it covers more than 100 languages, reads up to 8,192 tokens per chunk and needs no instruction prefix. Open WebUI's docs suggest nomic-embed-text for Ollama setups, and it is smaller at 274 MB, but its model card requires search_document: and search_query: prefixes on every text, and Open WebUI adds those only through environment variables.
-
Pull the embedding model into the template's Ollama over SSH:
ssh -i ~/.ssh/your_key ubuntu@<instance-ip> C=$(sudo docker ps -q --filter publish=8080) sudo docker exec "$C" ollama pull bge-m3 -
In Open WebUI, open Settings > Admin > Documents.
-
Set Embedding Model Engine to Ollama, leave the API Base URL it fills in, and enter
bge-m3as the Embedding Model. -
Set Text Splitter to Token (Tiktoken), Chunk Size to 1500, Chunk Overlap to 200 and Top K to 10, then click Save.
-
If you uploaded documents before this change, click Reindex under Reindex Knowledge and Memory Vectors.
The settings that matter#
The defaults are cautious, sized for small models with short contexts. These are the Open WebUI v0.11.4 defaults and what I change on a GPU with room for a 32K context.
| Setting | Default | What I set | Why |
|---|---|---|---|
| Embedding Model Engine | Default (SentenceTransformers) | Ollama | Embeds on the GPU |
| Embedding Model | sentence-transformers/all-MiniLM-L6-v2 | bge-m3 | Reads 8,192 tokens per chunk instead of 256 word pieces |
| Text Splitter | Default (Character) | Token (Tiktoken) | Chunk sizes count tokens, the unit the context is measured in |
| Chunk Size | 1000 | 1500 | Larger chunks keep more of each passage together |
| Chunk Overlap | 100 | 200 | Fewer sentences split across two chunks |
| Top K | 3 | 10 | More passages per question, for questions that span a document |
| Hybrid Search | Off | Off at first | Adds keyword matching, and a reranking model if you set one |
| Full Context Mode | Off | Off | Sends whole documents instead of chunks |
The chunk size, overlap and Top K values come from Open WebUI's own troubleshooting page, which suggests them for setups that mix local and cloud models. With the defaults, three chunks of 1,000 characters reach the model per question, 3,000 characters in all (our calculation: 3 x 1,000), which is little for a question that spans a long document.
The retrieved text has to fit in the chat model's context next to the conversation. At Top K 10 and 1,500-token chunks, one question can pull in up to 15,000 tokens (our calculation: 10 x 1,500), so I give the chat model a 32,768-token context. Ollama picks a default from GPU memory, 32,768 tokens from 24 to 48 GiB and 262,144 from 48 GiB up, capped at the model's maximum. Open WebUI's context field, though, pre-fills 2048 the moment you switch it on, which silently cuts retrieved text. Set it on purpose: in Settings > Admin > Models, edit the chat model and, under Advanced Params, set the context length (num_ctx) to 32768.
Hybrid Search adds keyword matching (BM25) to the vector search. Switched on, it shows a Reranking Model field that is empty by default, and without a model Open WebUI orders the combined results with the embedding model. A reranking model you enter on the default engine runs inside the Open WebUI process, so on the template it runs on the CPU too. I would leave Hybrid Search off until plain retrieval misses exact strings such as part numbers or error codes.
Build a knowledge base and ask it#
A knowledge base takes four steps to build and test.
- Open Workspace > Knowledge, click Create, and give the knowledge base a name and a description.
- Add files through Add Content: Upload files, Upload directory, or Sync directory to mirror a folder from your laptop. Wait until each file has finished processing.
- Start a new chat with your chat model, type
#, and pick the knowledge base from the list above the message box. - Ask a question whose answer you know is in the files, and check the citations under the answer against the source.
Attaching the knowledge base to a model in Workspace > Models works differently. Native function calling is Open WebUI's default, and in that mode knowledge attached to a model is not added to the prompt: the model receives knowledge tools and has to decide to call them. That suits models that call tools reliably. With a local model I would start with # in the chat, which puts the retrieved chunks in the prompt directly, and attach knowledge to a model only after it passes the known-fact test that way.
If a file answers nothing, open it in the knowledge base and look at the extracted text. A scanned PDF has no text layer, and Open WebUI's docs point to OCR-capable extraction engines such as Apache Tika or Docling for those.
GPU memory for the embedding model and the chat model#
Both models share the GPU, and nearly all of the memory they use belongs to the chat model and its context. Ollama keeps up to three models loaded per GPU by default and unloads each one 5 minutes after its last request, so the embedding model stays in memory while you upload and ask.
| Chat model | Chat model download | With bge-m3 (1.2 GB) | GPU I would use |
|---|---|---|---|
| gpt-oss:20b | 14 GB | 15.2 GB of weights | RTX A6000, 48 GB |
| gemma4:31b | 20 GB | 21.2 GB of weights | RTX A6000, 48 GB |
| gpt-oss:120b | 65 GB | 66.2 GB of weights | H200 NVL, 141 GB |
The sizes are the Ollama library downloads on 2026-09-28, and the sums are our calculation. The context comes on top: every token of a 32K context has a KV cache entry in GPU memory, and the size per token depends on the model's attention layout, which the KV cache guide explains. ollama ps inside the app container shows what each loaded model really takes and whether any of it spilled to the CPU. The Ollama GPU requirements guide covers larger chat models.
Keep the knowledge base when the disk is deleted#
Stopping an instance terminates it and deletes its disk: the knowledge bases, their vectors, your uploads, your chats and the models go with it. QuantaCloud has no volumes or snapshots, so a knowledge base survives a stop only if you copy it off first. There are three ways, from the simplest to the most complete.
The first is to keep your source files. If the documents live on your laptop or in a repository, the next instance needs only the settings above and an Upload directory or Sync directory of that folder. Embedding runs again, which is where the GPU engine pays off.
The second is Open WebUI's export. In Workspace > Knowledge, the menu on a knowledge base has Export, visible to admins, which downloads a zip with the extracted text of every file as a .txt file. On the next instance, unzip it and upload the folder into a new knowledge base. Retrieval works from that extracted text anyway, so the new knowledge base answers from the same words. What you lose are the original files and their formatting.
The third is to copy Open WebUI's data folder, which keeps accounts, chats, settings, knowledge bases and their vectors together in webui.db, uploads/ and vector_db/. Stop the container before you copy it, so no database is written to mid-copy, and start it again afterwards. On the template, this also stops and restarts the Ollama inside it:
C=$(sudo docker ps -q --filter publish=8080) # on the Compose setup: C=open-webui
sudo docker stop "$C"
sudo docker cp "$C":/app/backend/data ./open-webui-data
sudo tar czf open-webui-data.tgz open-webui-data
sudo docker start "$C"
Then download the archive from your laptop, and on the Compose setup the .env file too, so the same WEBUI_SECRET_KEY signs the logins:
scp -i ~/.ssh/your_key ubuntu@<instance-ip>:open-webui-data.tgz .
scp -i ~/.ssh/your_key ubuntu@<instance-ip>:open-webui/.env .
I would restore a data folder only on the Docker Compose setup from the setup guide, where you control the containers. On the new instance, create ~/open-webui with the same compose.yaml, and upload the archive and the .env file into it. An archive from the template comes without a .env file, so create a new one as the setup guide shows. Then pull the models and copy the data into the Open WebUI container before it starts for the first time, so the restored database never mixes with a fresh one:
cd ~/open-webui && tar xzf open-webui-data.tgz
sudo docker compose up -d ollama
sudo docker exec ollama ollama pull bge-m3
sudo docker exec ollama ollama pull gpt-oss:20b
sudo docker compose create open-webui
sudo docker cp ./open-webui-data/. open-webui:/app/backend/data/
sudo docker compose up -d
Restore into the same Open WebUI version or a newer one, since Open WebUI upgrades its database when it starts and does not downgrade it. Settings > About shows the version the archive came from, and the image tag in compose.yaml sets the version it goes into.
An archive from the template needs one more change once the stack is up. The template's Open WebUI reached Ollama at http://localhost:11434, and the restored settings keep that address, which does not work on the Compose setup. Set it to http://ollama:11434 in Settings > Admin > Connections for the chat models, and in the embedding API Base URL in Settings > Admin > Documents.
The file transfer guide covers moving larger archives, and the data persistence docs list what else lives on the disk.
When answers go wrong#
Most bad answers come from retrieval, not from the chat model.
| Symptom | Likely cause | Fix |
|---|---|---|
| The model says the documents do not mention it | Knowledge attached to a model that did not call its tools | Attach the knowledge base in the chat with # |
| Answers use only part of a long document | Top K too low, or the context cuts the retrieved text | Raise Top K and set the context length on the model |
| Nonsense matches after changing the embedding model | Old vectors from the previous model | Click Reindex, and upload chat files again |
| A file answers nothing | Scanned PDF with no text layer | Check the extracted text, then use an OCR-capable extraction engine |
| Uploads take long to process | The default embedding model running on the CPU | Switch the engine to Ollama with bge-m3 |
My rule: set the embedding engine to Ollama with bge-m3 before the first upload, attach knowledge with # until the model has proved it uses its tools, and give the chat model a context at least twice the retrieved text. Export the knowledge base, or keep the source folder, before you press Stop. For the chat side of the same template, the Open WebUI page covers pricing and privacy, and the setup guide has the Compose file.