The DeepSeek model that runs on one GPU is a distill, not DeepSeek-R1 itself. The R1 distills are Qwen and Llama models from 1.5B to 70B parameters, trained on R1's reasoning, and the 32B distill is 65.5 GB in BF16, which one 80 GB or 96 GB GPU holds, while a 48 GB card holds its 20 GB Ollama download with a 32k context. The full DeepSeek-R1 and V3 models have 671B parameters and ship as 688.6 GB of FP8 weights, which is more GPU memory than any single on-demand VM on QuantaCloud offers. They are a job for a reserved server, as is DeepSeek V4-Pro at 892.7 GB.
Which DeepSeek model you mean#
"DeepSeek" covers three kinds of model, and their memory needs differ by a factor of more than 200 (our calculation: 892.74 GB against 3.55 GB). These are the open-weight releases people self-host, with the file sizes on Hugging Face:
| Model | What it is | Parameters (total, active) | Weights on Hugging Face | Context | Licence |
|---|---|---|---|---|---|
| R1-Distill-Qwen 1.5B to 32B, R1-Distill-Llama 8B and 70B | Dense Qwen2.5 and Llama models trained on R1's outputs | 1.8B to 70.6B | 3.55 to 141.11 GB, BF16 | 131,072 | MIT, plus the base model's licence |
| R1-0528-Qwen3-8B | Qwen3-8B trained on R1-0528's reasoning | 8.2B | 16.38 GB, BF16 | 131,072 | MIT |
| R1, R1-0528 | Mixture-of-experts reasoning model | 671B, 37B | 688.59 GB, FP8 | 128K | MIT |
| V3, V3-0324, V3.1, V3.1-Terminus | The same 671B architecture as a chat model | 671B, 37B | 688.59 GB, FP8 | 128K | V3: DeepSeek Model License. Later: MIT |
| V3.2 | V3 with DeepSeek Sparse Attention | 671B, 37B | 689.48 GB, FP8 | 163,840 in its config | MIT |
| V4-Flash-0731 | Mixture of experts with FP4 experts | 284B, 13B | 166.89 GB | 1M | MIT |
| V4-Pro-0813 | Mixture of experts with FP4 experts | 1.6T, 49B | 892.74 GB | 1M | MIT |
| V4.1-Flash | Multimodal mixture of experts with Engram memory | 552B plus a 196B Engram, 8B to 16B | 510.30 GB | 1M | MIT |
A mixture-of-experts model needs all of its experts in memory, even though each token only uses a few of them, so memory follows the total parameters and the file size, never the active count. The VRAM guide explains the general rule.
The R1 distills fit one GPU#
Every R1 distill up to 32B fits a single QuantaCloud GPU in full BF16 precision. The table uses vLLM's default of 92% of the GPU and a 32,768-token context (our calculation), and the Ollama column is the library's default 4-bit tag:
| Distill | Base model | BF16 weights | KV cache per token | Smallest GPU for BF16 at 32k | Ollama tag |
|---|---|---|---|---|---|
| R1-Distill-Qwen-1.5B | Qwen2.5-Math-1.5B | 3.55 GB | 28 KiB | Any 48 GB card | deepseek-r1:1.5b, 1.1 GB |
| R1-Distill-Qwen-7B | Qwen2.5-Math-7B | 15.23 GB | 56 KiB | Any 48 GB card | deepseek-r1:7b, 4.7 GB |
| R1-Distill-Llama-8B | Llama-3.1-8B | 16.06 GB | 128 KiB | Any 48 GB card | deepseek-r1:8b-llama-distill-q4_K_M, 4.9 GB |
| R1-0528-Qwen3-8B | Qwen3-8B | 16.38 GB | 144 KiB | Any 48 GB card | deepseek-r1:8b, 5.2 GB |
| R1-Distill-Qwen-14B | Qwen2.5-14B | 29.54 GB | 192 KiB | Any 48 GB card | deepseek-r1:14b, 9.0 GB |
| R1-Distill-Qwen-32B | Qwen2.5-32B | 65.53 GB | 256 KiB | 80 GB, or 96 GB for more context | deepseek-r1:32b, 20 GB |
| R1-Distill-Llama-70B | Llama-3.3-70B-Instruct | 141.11 GB | 320 KiB | Two GPUs, see below | deepseek-r1:70b, 43 GB |
The 32B distill is the one I would start with. In BF16 it leaves room for about 51,000 tokens of cache on an 80 GB A100 or H100 PCIe and more than twice that on the 96 GB RTX PRO 6000 (our calculation). Its FP8 checkpoint, RedHatAI's per-channel build at 34.33 GB, fits a 48 GB L40S, L40 or RTX 6000 Ada, but a 32,768-token context, which takes 8 GiB, is borderline there: about 9.4 GiB is left before activations with ECC on. Those Ada cards run it with FP8 math, while the RTX A6000 and A100 run FP8 checkpoints weight-only, as the vLLM quantization guide explains.
Two Ollama details matter here. ollama pull deepseek-r1 with no tag gets the 8B distill of R1-0528, a 5.2 GB model, not the 671B one. And every distill tag carries a 131,072-token context, so on a large GPU Ollama's default context makes deepseek-r1:32b need about 34 GB of cache on top of its 20 GB of weights, more than a 48 GB card holds. Set num_ctx to 32768 and it needs 8.6 GB. The Ollama GPU requirements explain the default.
The 70B distill needs two GPUs, or FP8#
R1-Distill-Llama-70B is 141.1 GB in BF16, which misses even one 141 GB H200 NVL once vLLM keeps its 8% margin. Four setups hold it (our calculation):
| Setup | Weights | Left for the KV cache |
|---|---|---|
| FP8 checkpoint on one H200 NVL | 72.67 GB | About 61.5 GiB, roughly 200,000 tokens |
| FP8 checkpoint on one RTX PRO 6000 | 72.67 GB | About 20.3 GiB, roughly 66,000 tokens |
BF16 on 2x H200 NVL, --tensor-parallel-size 2 | 141.11 GB | About 127 GiB |
BF16 on 4x 48 GB, --pipeline-parallel-size 4 | 141.11 GB | About 34 GiB |
The FP8 build on one H200 NVL is the simplest of the four. The vLLM multi-GPU guide covers the two-GPU and four-GPU setups. With Ollama, the 43 GB deepseek-r1:70b fits a 48 GB card only with a short context such as 8,192 tokens, and an 80 GB card with room to spare at 32,768 (our calculation).
The full 671B models need more than one VM holds#
DeepSeek-R1, R1-0528 and the V3 family ship as 688.59 GB of FP8 weights, so they need about 700 GB of GPU memory before any cache. The largest on-demand VM in the 2026-09-27 catalog, 8x A100 SXM4, has 640 GB, and the A100 has no FP8 math anyway. The cache itself is small for a model this size: R1's multi-head latent attention stores 70,272 bytes per token, 9.2 GB for a 131,072-token conversation (our calculation). vLLM keeps a full copy of that latent cache on every GPU under tensor parallelism, and its docs suggest data parallelism for the attention layers of models that use this attention, DeepSeek among them.
For the real thing, an 8-GPU H200 server has 1,128 GB and an 8-GPU B200 server 1,440 GB (our calculation from NVIDIA's per-GPU figures), and both hold the FP8 weights with room for many conversations. We build reserved servers like these to order: send a capacity brief with the model and the traffic you expect, and the configuration, lead time and terms come back in writing. B200 servers and GPU clusters describe the hardware.
Community quantizations shrink R1 enough for a multi-GPU VM, at a cost in quality. These GGUF files run on llama.cpp, which splits them by layers across the GPUs in one VM. The right-hand column compares each file with the memory nvidia-smi reports for the 2026-09-27 configurations, with the file sizes converted to GiB (our calculation):
| GGUF file | Size | Smallest on-demand VMs with room to spare |
|---|---|---|
| Unsloth R1 UD-IQ1_S | 140.2 GB | 1x H200 NVL, 140.4 GiB, or 2x A100 PCIe or 2x H100 PCIe, about 160 GiB |
| Unsloth R1 UD-Q2_K_XL | 226.6 GB | 2x H200 NVL, 280.8 GiB |
| Unsloth R1 Q3_K_M | 319.2 GB | 4x A100 PCIe, 320 GiB |
| Unsloth R1 Q4_K_M, or Ollama's deepseek-r1:671b | 404.4 GB and 404 GB | 8x A100 SXM4, 640 GiB |
The 4x A100 PCIe VM reports 320 GiB, which holds the Q3_K_M file, 297.3 GiB, with about 23 GiB to spare for the cache and llama.cpp's working buffers.
The cost grows as the bits fall: for a plain Q2_K file of an 8B model, llama.cpp's own figures show 20 times the perplexity increase of Q4_K_M. I would treat a 1 or 2-bit R1 as a way to try the model, not to judge it, and use a distill or a reserved server for real work. The llama.cpp guide covers the multi-GPU flags.
DeepSeek V4#
DeepSeek's newest open models use FP4 for their expert weights and FP8 for the rest, which makes them far smaller than their parameter counts suggest. V4-Flash-0731, with 284B parameters, is 166.89 GB, so on size alone it fits a 192 GB VM such as 2x RTX PRO 6000 or 4x 48 GB, and 2x H200 NVL, at 282 GB, leaves more room (our calculation). Its serving story is less settled than its size. DeepSeek's own vLLM example runs it on a single four-GPU GB300 node, the release ships no Jinja chat template, only Python scripts to encode messages, and its FP4 expert weights need matching kernels in the serving engine, so check vLLM's recipe for your GPU before you plan on it. V4-Pro-0813, at 892.74 GB, needs more GPU memory than any on-demand VM has, so it is a reserved-server model. V4.1-Flash, at 510.30 GB, fits the 640 GB of an 8x A100 SXM4 on size alone, but the A100 has neither FP8 nor FP4 math and the model card gives no vLLM command, so I would plan it on a reserved server as well. Ollama lists V4-Pro and V4.1-Flash only as cloud models, which run on Ollama's servers rather than on your GPU.
Licences#
DeepSeek's own weights are MIT-licensed for R1, R1-0528, V3-0324 and every later release, which allows commercial use, modification and distillation. The original V3 from December 2024 uses the DeepSeek Model License, which also supports commercial use. The distills carry a second licence from their base model, as DeepSeek's model card points out: the Qwen distills come from Qwen2.5, originally under Apache 2.0, R1-Distill-Llama-8B from Llama 3.1 under the Llama 3.1 licence, and R1-Distill-Llama-70B from Llama 3.3 under the Llama 3.3 licence. The Llama licences add conditions of their own, such as attribution and a separate licence from Meta above 700 million monthly active users, so read them before building a product on those two.
Run a distill on QuantaCloud#
The quickest route is the Open WebUI + Ollama template on an RTX A6000:
Launch Open WebUI + Ollama on an RTX A6000- Open the application from the Deployments page and create the Open WebUI admin account.
- Type
deepseek-r1:32binto the model selector and pull it, a 20 GB download. - In Chat Controls > Advanced Parameters, or in the model's settings, switch on
num_ctxand set it to 32768, so the model and its cache stay on the GPU. The control starts at 2048 when you switch it on, which is too short for R1's long reasoning.
For an API with many users, run vLLM from the vLLM Docker guide on the Bare Metal template, with VLLM_TAG and VLLM_API_KEY set as that guide shows. On an H200 NVL, the FP8 70B distill takes one command:
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
"vllm/vllm-openai:$VLLM_TAG" \
RedHatAI/DeepSeek-R1-Distill-Llama-70B-FP8-dynamic \
--max-model-len 32768 \
--reasoning-parser deepseek_r1 \
--api-key "$VLLM_API_KEY"
--reasoning-parser deepseek_r1 returns the model's thinking separately from its answer. DeepSeek's model card asks for a temperature between 0.5 and 0.7, with 0.6 recommended, and for all instructions in the user message rather than a system prompt, while R1-0528 and its Qwen3-8B distill accept system prompts. Stopping the instance deletes its disk, so every launch downloads the 72.67 GB of weights again.
DeepSeek GPU FAQ#
Can I run DeepSeek-R1 on one GPU?
Not the full 671B model: its FP8 weights are 688.6 GB, almost five times the 141 GB of the largest single GPU on QuantaCloud. The distills run on one GPU, from the 1.5B on any card to the 70B in FP8 on one H200 NVL.
How much VRAM does deepseek-r1:14b or 32b need?
In Ollama, deepseek-r1:14b is a 9.0 GB download and deepseek-r1:32b 20 GB, plus the KV cache for your context: 192 KiB and 256 KiB per token. At 32,768 tokens both fit a 48 GB card. In vLLM at BF16, the 14B needs 29.5 GB of weights and the 32B 65.5 GB.
What GPUs does DeepSeek use?
DeepSeek reports that training V3 took 2.788 million H800 GPU hours. To run its models, the size of the checkpoint decides: one GPU for the distills, and an 8-GPU H200 or B200 class server for the full 671B models in FP8.
Is DeepSeek free for commercial use?
R1, V3-0324 and later, and V4 are MIT-licensed, which allows commercial use. The Llama-based distills also carry Meta's Llama 3.1 or 3.3 licence, and the Qwen-based ones come from Apache 2.0 models.
The rule I follow for DeepSeek: run a distill on one GPU, sized by its file plus the context you need, and treat the full 671B models as a reserved-server project. For most people that means deepseek-r1:32b on a 48 GB RTX A6000 with num_ctx at 32768, or the FP8 70B distill on an H200 NVL when quality matters more than cost. The GPU catalog lists every GPU with its live price, and self-hosted LLM inference covers the serving side.