GPU guide

vLLM vs Ollama: which one to run on a GPU server

vLLM vs Ollama on one GPU: design, GGUF vs safetensors, batching and concurrency, memory use and OpenAI compatibility, and when to use each.

Faiz Ahmed13 min read

The short answer is: Ollama for one person or a small team chatting with a model, vLLM for an application that sends many requests at once. Ollama is built for convenience. It installs with one line, pulls ready-made model files that default to 4-bit builds, loads a model when a request names it and unloads it after five idle minutes, and by default works on one request per model at a time. vLLM is built for throughput. It serves Hugging Face checkpoints as they are published, batches every running request into each step on the GPU, and claims 92% of GPU memory at startup to hold as many requests as it can. Both speak the OpenAI API, so moving from one to the other is mostly a change of base URL.

vLLM and Ollama side by side#

The defaults below are from the two projects' current releases, Ollama 0.34.4 and vLLM 0.30.0.

Ollama 0.34.4vLLM 0.30.0
Built forChat for one person or a few, many models on demandMany concurrent requests to one model
Enginellama.cpp, per Ollama's READMEIts own engine, built on PagedAttention and continuous batching
Model filesGGUF from the Ollama library, or your own through a ModelfileHugging Face safetensors, with AWQ, GPTQ, FP8, NVFP4 or MXFP4 read from the checkpoint
Qwen3-8B as you would first run itqwen3:8b: Q4_K_M, 5.2 GBQwen/Qwen3-8B: BF16, 16.4 GB
Requests per model at once1 (OLLAMA_NUM_PARALLEL), with up to 512 more queuedUp to 256 on 48 GB cards and the A100, 1,024 on the H100, H200 NVL and RTX PRO 6000 (--max-num-seqs)
GPU memoryWeights plus a KV cache sized for context x parallel requests, per loaded model92% of the GPU at startup (--gpu-memory-utilization), shared by all requests
Models per serverUp to 3 per GPU, loaded on demandOne, plus LoRA adapters
Model too big for the GPUSplits layers between the GPU and system RAMRefuses to start, unless you offload weights with --cpu-offload-gb
OpenAI APIA subset of /v1/v1 chat, completions, responses, embeddings and models
AuthenticationNone on the local API--api-key, which guards only the /v1-style routes
Default address127.0.0.1, port 11434Every interface, port 8000
On QuantaCloudThe Open WebUI + Ollama template, or Bare MetalBare Metal with Docker or pip. There is no vLLM template

How each one is built#

Ollama is a model runner with a model library attached. Its README lists llama.cpp as its backend, and its Go server pulls models by name (ollama pull qwen3:8b), stores them, loads one when a request names it, and unloads it after five idle minutes (OLLAMA_KEEP_ALIVE). A Modelfile packages a model with its prompt template and parameters. The design goal is one command from nothing to a working chat.

vLLM is a serving engine built around one idea from its paper: keep the KV cache in fixed-size blocks, the way an operating system pages virtual memory, so almost no memory is wasted and more requests fit on the GPU at once. The authors report 2 to 4 times the throughput of FasterTransformer and Orca at the same latency. On top of that sits a scheduler that adds and removes requests between steps instead of waiting for a whole batch to finish, the iteration-level scheduling that the Orca paper introduced. vLLM is aimed at a GPU and a model you pick from Hugging Face, and by default it claims 92% of that GPU for itself.

GGUF vs safetensors#

The file format decides what you are really comparing. Ollama's library ships GGUF files, the single-file format of llama.cpp, and a model's default tag is a 4-bit Q4_K_M build: that holds for qwen3:8b (5.2 GB), llama3.1:8b, gemma4:31b and qwen3.6:27b alike. The same library has qwen3:8b-q8_0 at 8.9 GB and qwen3:8b-fp16 at 16 GB. You can bring your own GGUF file or a safetensors directory with a Modelfile (FROM /path/to/model) and ollama create, but Ollama does not quantize GGUF files during import.

vLLM loads safetensors checkpoints from Hugging Face and reads the format from each checkpoint's quantization_config. Qwen/Qwen3-8B runs in BF16 at 16.4 GB, and an AWQ, GPTQ, FP8 or NVFP4 build of the same model runs in that format with no extra flag. GGUF is not built in: it needs the vllm-gguf-plugin, and vLLM's own docs call that support "highly experimental and under-optimized". The vLLM quantization guide covers which format suits which GPU.

That is the trap in a quick side-by-side test: ollama run qwen3:8b and vllm serve Qwen/Qwen3-8B run the same model at different precision. The Ollama file is about a third of the size (5.2 GB against 16.4 GB, our calculation), and generating tokens one at a time is memory-bound, so the 4-bit file has less to read for every token. Compare like with like: qwen3:8b-fp16 against the BF16 checkpoint, or a 4-bit GGUF against an AWQ checkpoint.

Batching and concurrency#

Concurrency is where the two differ most. vLLM batches continuously. At every step its scheduler generates one token for each request already running, fills the rest of the step's token budget (--max-num-batched-tokens) with prompt tokens from waiting requests, chunking long prompts, and admits new requests as memory frees up. Up to --max-num-seqs requests run together: 256 by default on 48 GB cards and the A100, and 1,024 on GPUs with 70 GiB or more, such as the H100, H200 NVL and RTX PRO 6000. When the KV cache fills, vLLM preempts a request and recomputes it once space frees, rather than failing it.

Ollama works on OLLAMA_NUM_PARALLEL requests per model at a time, and the default is 1. Everything else waits in a queue of up to 512 requests (OLLAMA_MAX_QUEUE), and past that Ollama answers with a 503. You can raise the parallel count, but every slot gets a full context window of its own. Ollama's FAQ puts it this way: a 2K context with 4 parallel requests becomes an 8K context, with the memory to match.

The arithmetic for 32 users shows the difference (our calculation). Qwen3-8B keeps 144 KiB of KV cache per token. Thirty-two Ollama slots of 32,768 tokens reserve 32 x 32,768 x 144 KiB = 144 GiB, three times the memory of a 48 GB card. vLLM keeps one shared pool instead: on the same card with ECC on, its 92% budget leaves about 26.1 GiB after the weights, room for about 190,000 tokens, and 32 requests of 1,280 tokens use about 41,000 of them. The KV cache guide has the formula for any model.

When the prompts are a file rather than live traffic, vLLM's LLM class runs the same batching with no server at all, as batch inference with vLLM shows.

Memory behaviour#

vLLM takes its memory up front. At startup it claims 92% of the GPU, loads the weights, measures what activations and CUDA graphs need, and turns the rest into KV cache blocks. nvidia-smi then shows the GPU almost full whether one request is running or none, and a second vLLM process on the same GPU fails to start unless both lower --gpu-memory-utilization. If the weights plus one full-length request do not fit, vLLM refuses to start and prints the reason, which fixing vLLM out-of-memory errors goes through error by error.

Ollama takes memory when a model is asked for, and gives it back. A model loads on its first request and stays for 5 minutes after the last one, and up to three models per GPU can be resident if they fit. When a model does not fit in VRAM, Ollama keeps part of it in system RAM, and ollama ps shows the split, for example 48%/52% CPU/GPU. That keeps a model running, slowly, where vLLM would stop. The KV cache is FP16 by default, and OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves it when flash attention is on.

Ollama also picks the default context from total VRAM: 4,096 tokens below 23 GiB, 32,768 from 23 GiB and 262,144 from 47 GiB, capped at the model's own maximum. Those thresholds apply to the memory the GPU reports, so a 48 GB card that reports at least 47 GiB gives Qwen3-8B its full 40,960 tokens, while an L40S, which reports 46,068 MiB (about 45 GiB) in public nvidia-smi output, gets 32,768. vLLM takes the maximum from the model config too, unless you set --max-model-len.

OpenAI compatibility#

Both serve the OpenAI API under /v1, with different gaps. vLLM implements /v1/chat/completions, /v1/completions, /v1/responses, /v1/embeddings and /v1/models, plus an Anthropic-style /v1/messages. Its chat requests accept logprobs, n and tool_choice, and structured output through response_format or vLLM's own structured_outputs field.

Ollama describes its /v1 as a subset of the OpenAI API. Chat and completions work without logprobs, n or tool_choice, images go in as base64 only, and /v1/responses is stateless. The OpenAI API has no field for context length, so on Ollama you set it on the server or bake num_ctx into a Modelfile.

Client code barely changes. The OpenAI SDK needs a base URL and a key: http://127.0.0.1:8000/v1 with your --api-key for vLLM, and http://127.0.0.1:11434/v1/ with any string for Ollama, which ignores it.

Running each on QuantaCloud#

The Open WebUI + Ollama template is the quickest Ollama path. Ollama runs inside the app container next to Open WebUI, the VM publishes only Open WebUI's port on 127.0.0.1, and the app URL sits behind your QuantaCloud login plus Open WebUI's own sign-in. Open WebUI on QuantaCloud covers the template, and the Ollama API on a remote GPU covers a plain Ollama server on the Bare Metal template.

Launch Open WebUI + Ollama on an RTX A6000

vLLM runs on the Bare Metal template, a plain Ubuntu 22.04 VM with the NVIDIA driver and Docker. QuantaCloud has no vLLM template and no managed endpoint, so you run the official image yourself, as in deploying vLLM with Docker, or install it with pip or uv. Open WebUI can use a vLLM server as an OpenAI connection if you want a chat window on top.

Launch an RTX A6000 for vLLM

Neither server is safe to expose as installed. Ollama's local API has no authentication at all. vLLM listens on every interface unless you pass --host, and its API key leaves routes such as /invocations open. Keep both on 127.0.0.1 and reach them through an SSH tunnel. Stopping an instance terminates it and deletes its disk, so every launch downloads the model again: 5.2 GB for qwen3:8b, or the 8.7 GB vLLM image plus 16.4 GB of BF16 weights for vLLM.

How to compare them on one GPU#

The fair test is the same prompts, the same model at the same precision, and the same GPU. Use vllm bench serve as the load generator for both servers, because it can drive any OpenAI-compatible endpoint, and run the two servers one at a time, since vLLM claims 92% of the GPU.

  1. Launch an RTX A6000 on the Bare Metal template and start vLLM with Qwen/Qwen3-8B and --max-model-len 32768, as in the vLLM Docker guide.

  2. Run the load from a second container on the host network, so it reaches both servers on 127.0.0.1. This is the 32-request run. For one request at a time, use --max-concurrency 1 --num-prompts 32 --seed 1.

    Terminal
    docker run --rm --network host --entrypoint vllm \
      -e OPENAI_API_KEY="$VLLM_API_KEY" \
      -v ~/.cache/huggingface:/root/.cache/huggingface \
      "vllm/vllm-openai:$VLLM_TAG" bench serve \
      --backend openai --host 127.0.0.1 --port 8000 --model Qwen/Qwen3-8B \
      --dataset-name random --random-input-len 1024 --random-output-len 256 --ignore-eos \
      --max-concurrency 32 --num-prompts 256 --seed 32
    
  3. Stop vLLM with docker stop vllm, install Ollama with the pinned script from the Ollama API guide, and pull the 16-bit tag with ollama pull qwen3:8b-fp16.

  4. Send Ollama the same prompts with the same seeds, first at its defaults with --max-concurrency 1 --num-prompts 32 --seed 1. Only the port and the served name change, and the tokenizer still comes from Qwen/Qwen3-8B:

    Terminal
    docker run --rm --network host --entrypoint vllm \
      -v ~/.cache/huggingface:/root/.cache/huggingface \
      "vllm/vllm-openai:$VLLM_TAG" bench serve \
      --backend openai --host 127.0.0.1 --port 11434 \
      --model Qwen/Qwen3-8B --served-model-name qwen3:8b-fp16 \
      --dataset-name random --random-input-len 1024 --random-output-len 256 --ignore-eos \
      --max-concurrency 1 --num-prompts 32 --seed 1
    
  5. For the 32-request run, give Ollama 32 slots of 4,096 tokens each, which covers a 1,280-token request and reserves 18 GiB of KV cache (our calculation). Run sudo systemctl edit ollama, add the lines below, run sudo systemctl restart ollama, then repeat step 4 with --max-concurrency 32 --num-prompts 256 --seed 32.

    ini
    [Service]
    Environment="OLLAMA_NUM_PARALLEL=32"
    Environment="OLLAMA_CONTEXT_LENGTH=4096"
    

ignore_eos is not among the request fields Ollama's OpenAI API supports, so its replies can stop before 256 tokens. Read its mean output length next to the tokens per second. To compare memory, run nvidia-smi --query-gpu=memory.used --format=csv -lms 500 in a second terminal during each run.

vLLM vs Ollama FAQ#

Is vLLM faster than Ollama?

The honest answer depends on concurrency and precision. With many requests at once, the defaults decide it: vLLM batches up to 256 requests per step on a 48 GB card, while Ollama works on one request per model and queues the rest. With one request, precision matters more than the engine: Ollama's default qwen3:8b is a 4-bit file a third the size of the 16-bit checkpoint, and generating each token means reading the weights.

Can vLLM run GGUF models?

Only through a plugin. GGUF support moved out of vLLM into vllm-gguf-plugin, and vLLM's docs describe it as highly experimental and under-optimized. If a model only exists as GGUF, Ollama or llama.cpp is the natural server for it.

Can Ollama serve many users at once?

Yes, within the memory you give it. Set OLLAMA_NUM_PARALLEL above its default of 1 and each slot reserves a full context window, so memory grows with the slot count times the context length. Up to 512 further requests wait in the queue before Ollama returns a 503.

Can I use Open WebUI with vLLM?

Yes. Open WebUI connects to any OpenAI-compatible server: add vLLM's base URL and API key under Admin Settings, Connections, or set OPENAI_API_BASE_URL and OPENAI_API_KEY. From inside a container, Open WebUI's docs use http://host.docker.internal:8000/v1 to reach a vLLM server on the same host.

Which GPU should I rent for each?

For Ollama with 4-bit models, a 48 GB RTX A6000 holds even the qwen3:32b tag, a 20 GB download, with room for context, and Ollama GPU requirements covers bigger models. For vLLM, the same card serves an 8B model in BF16, or a 32B model as a 4-bit AWQ checkpoint. A 32B model in FP8 belongs on an Ada card such as the L40S, where vLLM computes in FP8, and in BF16 it just fits an 80 GB card with one 32k conversation, so for more room it needs the 96 GB RTX PRO 6000. The VRAM guide covers the rest.


My rule: Ollama when the users are people typing, vLLM when the user is a program sending requests in parallel. Start either one on an RTX A6000 at $0.48/GPU-hr, keep it on 127.0.0.1, and move to vLLM the day your Ollama queue starts to grow. If throughput under load is the whole point, vLLM vs SGLang is the next comparison, and self-hosted LLM inference covers the options together.

Launch Open WebUI + Ollama Launch a GPU for vLLM

Keep building

Choose your next step.