GPU guide

Ollama GPU requirements: which models fit 48, 80, 96 and 141 GB

How much GPU memory Ollama models need: library sizes by quantization, the KV cache of Ollama's default context, CPU offload and which GPU fits.

Faiz Ahmed11 min read

The rule is simple: the model's download size plus the KV cache for your context has to fit in GPU memory, or Ollama moves part of the model to the CPU and slows down. Ollama's default tags are 4-bit files, mostly Q4_K_M, of about 0.6 GB per billion parameters: qwen3:8b is 5.2 GB, qwen3:32b 20 GB and llama3.3:70b 43 GB. The cache is the part people miss. With 47 GiB or more of GPU memory, Ollama picks a 256k-token default context, capped at the model's own limit, and for deepseek-r1:32b that cap is 131,072 tokens and about 34 GB of cache on top of its 20 GB of weights (our calculation).

Memory by model and quantization#

Most models in the Ollama library come in several quantizations, and the tag you pull decides the size. These are the library's own figures for the default 4-bit tag, the 8-bit tag and the 16-bit tag:

ModelParametersDefault tag (q4_K_M)q8_016-bitContext window
qwen3:8b8.2B5.2 GB8.9 GB16 GB40K
qwen3:14b14.8B9.3 GB16 GB30 GB40K
qwen3:32b32.8B20 GB35 GB66 GB40K
gemma4:31b31.3B20 GB34 GB63 GB256K
deepseek-r1:32b32.8B20 GB35 GB66 GB128K
llama3.3:70b70.6B43 GB75 GB141 GB128K
qwen3:235b235B142 GB250 GB470 GB256K
gpt-oss:20b20.9B14 GB (MXFP4)Not offeredNot offered128K
gpt-oss:120b117B65 GB (MXFP4)Not offeredNot offered128K

The default tag is the one you get when you type a model name without a quantization, and at about 0.6 GB per billion parameters it is the one to size for (our calculation: 20 GB over 32.8B, 43 GB over 70.6B). The 8-bit tag is about 1.7 times larger and the 16-bit tag 3 to 3.3 times. gpt-oss ships in OpenAI's own 4-bit MXFP4 format only.

Ollama's q4_K_M and q8_0 tags are llama.cpp's GGUF quantization types, the same files the llama.cpp guide sizes, and llama.cpp's quantize tool measures their cost on Llama-3-8B: Q4_K_M raises perplexity by 0.1754 over the 16-bit model and Q8_0 by 0.0026. My reading is that q4_K_M is the right default, and q8_0 is worth its size when the GPU has the room after the context. Newer models also list mlx, nvfp4 and mxfp8 tags, and the library labels all of them MLX, the engine Ollama's blog introduces for Apple silicon. On an NVIDIA GPU I would stick to the q4_K_M, q8_0 and 16-bit tags that this page sizes.

Context length is the other half#

The KV cache holds every token of the conversation, and Ollama sizes it from the context length before the first reply. With no setting of your own, Ollama picks the default from the total GPU memory it detects, summed across GPUs, and caps it at the model's trained context:

Total GPU memoryDefault context (Ollama 0.34.4)
Under 23 GiB4,096 tokens
23 to 47 GiB32,768 tokens
47 GiB and more262,144 tokens

The docs round the thresholds to 24 and 48 GiB, and the code uses 23 and 47 to allow for small differences in the memory a card reports. Every QuantaCloud GPU is sold as 48 GB or more, but a 48 GB card with ECC on, as the L40 and L40S ship, reports 46,068 MiB, below the 47 GiB line, and only with ECC off does it report 49,140 MiB, above it. So check what yours reports with nvidia-smi --query-gpu=memory.total --format=csv: a card that shows less than 48,128 MiB, which is 47 GiB, starts at 32,768 tokens.

On a GPU where Ollama detects 47 GiB or more, each model starts at its full trained context. That is where the memory goes (our calculation, 16-bit cache):

Model tagDownloadDefault context at 47 GiB and upKV cache per tokenKV cache at that contextTotal before overhead
qwen3:8b5.2 GB40,960144 KiB6.0 GB11.2 GB
gpt-oss:20b14 GB131,07224 KiB3.2 GB17.2 GB
qwen3:32b20 GB40,960256 KiB10.7 GB30.7 GB
qwen3.6:27b18 GB262,14464 KiB17.2 GB35.2 GB
qwen3:30b19 GB262,14496 KiB25.8 GB44.8 GB
deepseek-r1:32b20 GB131,072256 KiB34.4 GB54.4 GB
gpt-oss:120b65 GB131,07236 KiB4.8 GB69.8 GB
llama3.3:70b43 GB131,072320 KiB43.0 GB85.9 GB
qwen3:235b142 GB262,144188 KiB50.5 GB192.5 GB

The per-token figure comes from the model's config: 2 x attention layers x KV heads x head size x 2 bytes, counting only the layers whose cache grows with context. That is why gpt-oss and qwen3.6 get long contexts cheaply, while deepseek-r1:32b, a Qwen2.5 model underneath, needs more memory for its cache than for its weights. The KV cache guide works the formula for more models, and how much VRAM you need covers other workloads.

You change the context in three places. OLLAMA_CONTEXT_LENGTH=32768 on the server sets the default for every model. "options": {"num_ctx": 32768} sets it for one API request, and PARAMETER num_ctx 32768 in a Modelfile bakes it into a model variant, which is the route for OpenAI-compatible clients because that API has no context field. At 32,768 tokens deepseek-r1:32b needs 8.6 GB of cache instead of 34.4 GB, and llama3.3:70b 10.7 GB instead of 43.0 GB.

Two settings multiply or divide the cache. Parallel requests multiply it: Ollama's FAQ says a 2K context with 4 parallel requests becomes an 8K context, so OLLAMA_NUM_PARALLEL=4 needs four times the cache. OLLAMA_KV_CACHE_TYPE=q8_0 halves it, and q4_0 cuts it to about a quarter, with Flash Attention, which Ollama turns on by itself where the GPU supports it. The setting applies to every model on the server. ollama ps shows the context Ollama actually chose, in its CONTEXT column. When many people share one model, vLLM vs Ollama explains where vLLM's batching fits better.

What happens when a model does not fit#

Ollama does not refuse a model that is too big. Version 0.34.4 hands the model to llama.cpp's llama-server with the context it chose, and llama.cpp puts as many layers on the GPU as fit and runs the rest on the CPU from system RAM. If the load itself runs out of memory and you did not set the context, Ollama retries once with a smaller automatic one: 32,768 tokens, or 4,096 if it was already at 32,768 or less. The PROCESSOR column of ollama ps shows the split: 100% GPU means everything is on the GPU, and 48%/52% CPU/GPU means about half the model runs on the CPU. Ollama's own docs say to avoid offloading to the CPU for performance, and a split is the first thing I look for when a model feels slow.

Offloaded layers need system RAM as well. The RTX A6000 1x offers came with 24 to 64 GB of RAM in the 2026-09-27 catalog, so an offloaded 70B model can run short of RAM too. The fix, in the order I would try it:

  1. Lower the context with num_ctx to what you actually use. For most chat, 32,768 tokens is plenty.
  2. Switch the cache to q8_0 with OLLAMA_KV_CACHE_TYPE.
  3. Pull a smaller quantization, or a smaller model.
  4. Move to the next GPU tier, or to a VM with two GPUs.

The num_gpu request option sets the number of layers on the GPU by hand, and it defaults to letting Ollama decide. I leave it alone and fix the context instead.

Which QuantaCloud GPU fits which model#

Each tier below is a different instance, and the template is the same on all of them. Fits means the model and its cache stay on the GPU at the context shown (our calculation from the tables above):

GPU memoryGPUsAt Ollama's default contextWith num_ctx 32768
48 GBRTX A6000, RTX 6000 Ada, L40, L40SUp to qwen3:32b, qwen3.6:27b and gpt-oss:20b. qwen3:30b is tightAdds deepseek-r1:32b and qwen3:32b-q8_0
80 GBA100 80GB, H100 PCIeAdds gpt-oss:120b and qwen3:32b in 16-bitAdds llama3.3:70b
96 GBRTX PRO 6000Adds llama3.3:70b at its full 131,072 tokensAdds llama3.3:70b-instruct-q8_0
141 GBH200 NVLAdds llama3.3:70b-instruct-q8_0 with its full contextRoom for two large models at once, such as gpt-oss:120b and qwen3:32b
192 to 640 GB2x RTX PRO 6000, 2x H200 NVL, 4x 48 GB, 8x A100qwen3:235b on 2x RTX PRO 6000 or 2x H200 NVLqwen3:235b on 2x RTX PRO 6000 or 4x 48 GB

On a 48 GB card that reports less than 47 GiB, Ollama's default is already 32,768 tokens, so read the right-hand column for it. With several GPUs in one VM, Ollama loads a model onto a single GPU when it fits there and spreads it across all of them only when it does not, and CUDA_VISIBLE_DEVICES limits which GPUs it uses. The largest DeepSeek tag, deepseek-r1:671b at 404 GB, fits with room to spare only on the 8x A100 SXM4; the DeepSeek GPU requirements cover that family. Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 6 Oct 2026, 02:35 UTC

Launch Open WebUI + Ollama on an RTX A6000

Ollama GPU FAQ#

How much VRAM do I need for Ollama?

For the default 4-bit tags, about 0.6 GB per billion parameters for the weights, plus the KV cache for your context. An 8B model fits almost any GPU, a 32B model needs about 20 GB plus its cache, and a 70B model needs 43 GB plus its cache. On a 48 GB card that means a 70B model only with a short context.

Does Ollama need a GPU?

No, it runs on the CPU too, only much more slowly. For NVIDIA GPUs it needs compute capability 5.0 or newer, which every QuantaCloud GPU has, and driver 550 or newer, which nvidia-smi shows.

Why does ollama ps show a CPU/GPU split?

The model plus its KV cache did not fit in GPU memory, so Ollama put some layers on the CPU. On a large GPU the usual cause is the default context, which starts at the model's full trained length. Lower num_ctx, or use a q8_0 cache.

Can Ollama use more than one GPU?

Yes. It keeps a model on one GPU when it fits and spreads it across all the GPUs when it does not. The default context is based on the memory of all the GPUs together.

How do I set up Ollama on a QuantaCloud GPU?

The Open WebUI + Ollama template starts both, with no models preinstalled. The Open WebUI setup guide walks through it, and the Ollama API guide covers a plain Ollama server on the Bare Metal template.


My rule for Ollama: size for the default tag plus the cache of the context you will actually use, then set num_ctx to that instead of trusting the default. Up to 32B models, a 48 GB RTX A6000 does the job. For llama3.3:70b with a long context, go to the 96 GB RTX PRO 6000, and past 141 GB, to two GPUs in one VM. Open WebUI on QuantaCloud describes the template that runs all of this, and self-hosted LLM inference compares Ollama with the other serving stacks.

Deploy Open WebUI + Ollama

Keep building

Choose your next step.