gpt-oss-120b needs one GPU with at least 80 GB of memory, and gpt-oss-20b needs about 16 GB. That is OpenAI's own sizing, and the files agree: the 120b weights are 65.2 GB and the 20b weights 13.8 GB, because both ship with their mixture-of-experts weights already in 4-bit MXFP4. On QuantaCloud, gpt-oss-120b fits a single H100 PCIe (80 GB), RTX PRO 6000 (96 GB) or H200 NVL (141 GB), and gpt-oss-20b fits a single RTX A6000 (48 GB) with room for its full context.
What each model needs#
| Model | Weights on disk | OpenAI's sizing | QuantaCloud GPUs that fit |
|---|---|---|---|
| gpt-oss-120b | 65.2 GB | "a single 80GB GPU" | H100 PCIe, RTX PRO 6000, H200 NVL |
| gpt-oss-20b | 13.8 GB | "within 16GB of memory" | Every on-demand GPU, 48 GB and up |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Prices checked 6 Oct 2026, 02:35 UTC
What gpt-oss-120b and gpt-oss-20b are#
gpt-oss is OpenAI's open-weight model family, released on 2025-08-05 and still its latest open-weight release in September 2026. Both models are mixture-of-experts reasoning models. Each token runs through 4 experts, but every expert has to sit in GPU memory, so memory follows the total parameter count, not the active one.
| gpt-oss-120b | gpt-oss-20b | |
|---|---|---|
| Total and active parameters | 116.83B and 5.13B | 20.91B and 3.61B |
| Experts (used per token) | 128 (4) | 32 (4) |
| Layers (full attention) | 36 (18) | 24 (12) |
| Context | 131,072 tokens | 131,072 tokens |
| Weights on disk | 65.25 GB | 13.76 GB |
| KV cache per token, BF16 | 36 KiB | 24 KiB |
Both use OpenAI's harmony response format, and both let you set the reasoning effort to low, medium or high in the system prompt ("Reasoning: high"). The licence is Apache-2.0, plus a two-sentence usage policy that asks you to comply with applicable law.
Where the memory goes#
The weights are the fixed cost. 116.83B parameters at half a byte each would be 58.4 GB, but the checkpoint is 65.25 GB, because attention, the router and the embeddings stay in BF16 and the MXFP4 blocks carry scales. Never upcast it: in BF16 the 120b weights would be 233.7 GB (our calculation). The Hugging Face repo also holds two more copies of the weights, in original/ and metal/, which bring it to 195.8 GB. vLLM skips original/ by default and downloads only the 65.2 GB it needs. If you fetch the files yourself (downloading Hugging Face models fast covers the hf CLI), exclude both folders:
hf download openai/gpt-oss-120b --exclude "original/*" --exclude "metal/*"
The KV cache is small for a model this size. Half of gpt-oss's layers use a 128-token sliding window, so only the 18 full-attention layers of the 120b grow with context. That is 2 x 18 layers x 8 KV heads x 64 x 2 bytes, or 36,864 bytes per token, and a full 131,072-token conversation needs 4.5 GiB (our calculation). vLLM takes 92% of the card by default, and what is left after the weights decides how many long conversations fit at once:
| GPU | vLLM budget at 0.92 | Left after the weights | Full 131,072-token conversations |
|---|---|---|---|
| H100 PCIe, 80 GB (120b) | 73.3 GiB | 12.5 GiB | about 2.8 |
| RTX PRO 6000, 96 GB (120b) | 87.9 GiB | 27.2 GiB | about 6.0 |
| H200 NVL, 141 GB (120b) | 129.2 GiB | 68.4 GiB | about 15.2 |
| RTX A6000, 48 GB (20b) | 41.4 GiB | 28.6 GiB | about 9.5 |
These are upper bounds from our calculation, using the memory each card reports in nvidia-smi and ignoring activations and CUDA graphs. The RTX A6000 row assumes ECC is on, and with ECC off the card reports 49,140 MiB, which adds 2.8 GiB. The KV cache guide explains the formula.
Which GPUs can run gpt-oss#
vLLM's MXFP4 path needs compute capability 8.0 or newer, and every GPU QuantaCloud offers on demand meets that. Meeting it is not the same as being supported. vLLM's gpt-oss recipe lists the H100, H200 and B200 as verified and runs the A100 through a Triton attention backend and Marlin MXFP4 kernels, while Ampere, Ada and RTX 5090-class Blackwell are still marked as work in progress. The RTX A6000 is Ampere, the L40S is Ada, and the RTX PRO 6000 has the same compute capability as the RTX 5090 (12.0).
| GPU | Memory | Compute capability | gpt-oss-120b | gpt-oss-20b |
|---|---|---|---|---|
| RTX A6000 | 48 GB | 8.6 | Does not fit | Fits |
| L40S, L40, RTX 6000 Ada | 48 GB | 8.9 | Does not fit | Fits |
| A100 80 GB | 80 GB | 8.0 | Fits | Fits |
| H100 PCIe | 80 GB | 9.0 | Fits, tight | Fits |
| RTX PRO 6000 | 96 GB | 12.0 | Fits | Fits |
| H200 NVL | 141 GB | 9.0 | Fits | Fits |
Run gpt-oss-120b with vLLM#
The setup is the one in the vLLM Docker guide: the Bare Metal template, the driver check that picks v0.30.0 (R580 or newer) or v0.30.0-cu129, the GPU check, and an API key in VLLM_API_KEY. Then:
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v vllm-cache:/root/.cache/vllm \
"vllm/vllm-openai:$VLLM_TAG" \
openai/gpt-oss-120b \
--max-model-len 131072 \
--api-key "$VLLM_API_KEY"
For gpt-oss-20b, change the model to openai/gpt-oss-20b. Neither repo is gated, so you need no Hugging Face token. On the H100 PCIe the fit is tight: if startup runs out of memory, vLLM's recipe suggests adding --gpu-memory-utilization 0.95 --max-num-batched-tokens 1024. For function calling, add --tool-call-parser openai --enable-auto-tool-choice. Then send a request:
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $VLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-oss-120b", "messages": [{"role": "system", "content": "Reasoning: low"}, {"role": "user", "content": "Why does a mixture-of-experts model need all experts in memory?"}], "max_tokens": 500}'
Chat Completions returns the reasoning separately from the final answer. OpenAI recommends the Responses API for this model, and vLLM serves it at /v1/responses. Keep port 8000 on 127.0.0.1: the Docker guide shows the SSH tunnel and a reverse proxy for reaching it from elsewhere. Stopping the instance deletes its disk, so every new launch downloads the 65.2 GB of weights again.
Run gpt-oss-20b with Ollama on the Open WebUI template#
The shortest path to chatting with gpt-oss-20b is the Open WebUI + Ollama template, with no terminal at all.
Launch Open WebUI + Ollama on an RTX A6000- When the deployment is running, click Open Application (templates in the docs). The URL sits behind your QuantaCloud login, and Open WebUI then asks you to create its own account. The first account on a fresh install becomes the admin.
- Type
gpt-oss:20bin the model selector and pull it, a 14 GB download.gpt-oss:120bis 65 GB and needs an 80 GB or larger GPU. Skip the-cloudtags, which are Ollama's hosted models and never download to your GPU. - Start a chat.
Ollama picks the default context from total VRAM and caps it at the model's own maximum: 4k below 23 GiB, 32k from 23 GiB and 256k from 47 GiB (its docs round these to 24 and 48). On the RTX A6000 that gives gpt-oss-20b its full 131,072 tokens only with ECC off, when the card reports 49,140 MiB, and 32,768 with ECC on, when it reports 46,068 MiB, about 45 GiB. Check what it chose over SSH, finding the template's container by the port it publishes:
sudo docker exec -it $(sudo docker ps -q --filter publish=8080) ollama ps
The pulled model lives on the instance's disk, so stopping the instance deletes it and the next launch pulls it again. Open WebUI on QuantaCloud and the Open WebUI with Ollama setup guide cover the template in detail.
H100 PCIe, RTX PRO 6000 or H200 NVL for gpt-oss-120b#
The core problem with this choice is that the three cards win on different things. Memory and price favour the RTX PRO 6000. On 2026-09-27 it listed at $2.39 an hour against $2.59 for the H100 PCIe (live now: $2.39/GPU-hr and $2.59/GPU-hr), so the card with 16 GB more memory was 20 cents an hour cheaper. Memory bandwidth favours the H200 NVL at 4.8 TB/s, against 2.0 TB/s on the H100 PCIe and 1.6 to 1.8 TB/s on the RTX PRO 6000 depending on the edition. At small batch sizes, generating tokens is mostly a matter of reading weights from memory, so I expect the H200 NVL to lead on single-request speed.
For one user or a small team, I would start on the RTX PRO 6000, provided vLLM runs gpt-oss on it: it holds the model with room for about six full-length conversations at the lower price. For team chat with many people at once, the H200 NVL has the bandwidth and room for about 15 full-length conversations, and the cost per million tokens decides whether that is worth its higher hourly rate. For long-context RAG, where every request carries tens of thousands of tokens, the H200 NVL leaves 68.4 GiB for the cache against 27.2 GiB on the RTX PRO 6000 (our calculation). I would take the H100 PCIe only when the other two are unavailable, because it holds the model with the least room to spare. The RTX PRO 6000 vs H100 and RTX PRO 6000 vs H200 comparisons go beyond gpt-oss.
Fine-tuning gpt-oss#
Unsloth publishes the minimums: QLoRA needs 14 GB for gpt-oss-20b and 65 GB for gpt-oss-120b, and 16-bit LoRA needs 44 GB and 210 GB. OpenAI says the 120b can be fine-tuned on a single H100 node and the 20b even on consumer hardware. On QuantaCloud, that makes a 48 GB RTX A6000 enough for QLoRA on the 20b and just enough for its 16-bit LoRA at short sequence lengths. For QLoRA on the 120b, an 80 GB card is the floor, and the 96 GB RTX PRO 6000 or the 141 GB H200 NVL leaves room for longer sequences. 16-bit LoRA on the 120b needs a multi-GPU VM such as 2x H200 NVL (282 GB) or 4x RTX PRO 6000 (384 GB). The LoRA, QLoRA and full fine-tuning VRAM guide and the Unsloth fine-tuning tutorial go further.
gpt-oss hardware FAQ#
Can gpt-oss-120b run on a 48 GB GPU?
No. The weights alone are 65.2 GB, more than a 48 GB card holds. Two 48 GB cards hold it on paper: split across the pair with --tensor-parallel-size 2, or --pipeline-parallel-size 2 on cards without NVLink, the 60.8 GiB of weights leave about 22.0 GiB for the cache (our calculation, with ECC on).
vLLM still lists Ampere and Ada support for gpt-oss as work in progress, so I would rent one 96 GB RTX PRO 6000 instead.
Does gpt-oss-120b need an H100?
No. OpenAI names the H100 as an example of a single 80 GB GPU. Any GPU with 80 GB or more and compute capability 8.0 or newer meets the requirement, and on QuantaCloud that means the H100 PCIe, the A100 80 GB, the RTX PRO 6000 and the H200 NVL.
What context length can gpt-oss use?
131,072 tokens for both models, and vLLM uses the full length by default. A full-length conversation costs 4.5 GiB of KV cache on the 120b and 3 GiB on the 20b (our calculation), so on these cards the number of simultaneous conversations runs out long before the context length does.
How much VRAM does gpt-oss-20b need?
OpenAI says it runs within 16 GB. The weights are 13.8 GB, so a 16 GB card leaves little room for context, while a 48 GB card holds the weights plus more than nine full-length conversations of cache (our calculation).
Is gpt-oss free for commercial use?
The licence is Apache-2.0, which allows commercial use, plus OpenAI's usage policy, which asks you to comply with applicable law.
The rule I follow for gpt-oss is to run the 20b on the lowest-priced 48 GB card and the 120b on the smallest card that leaves room for the context and concurrency you need. For a small team that is the RTX PRO 6000. For many concurrent users or long documents, it is the H200 NVL.
Launch an RTX PRO 6000 for gpt-oss-120b Launch an H200 NVL for gpt-oss-120b