gpt-oss-20b fine-tunes with LoRA on one 80 GB GPU using the recipe in OpenAI's cookbook: Transformers unpacks the model's 4-bit MXFP4 experts into BF16, PEFT attaches LoRA adapters to the attention layers and to the experts of three layers, and TRL trains on conversations in gpt-oss's harmony format. OpenAI's cookbook reports about 18 minutes on an H100 for its 1,000-example dataset. On QuantaCloud, the lowest-priced single 80 GB GPU in the 2026-09-27 catalog was the A100 SXM4 at $1.50 an hour, now $1.50/GPU-hr. The script below is that recipe updated for the current TRL 1.14, PEFT 0.21 and Transformers 5.17, followed by merging the adapter and serving it with vLLM.
How much memory each recipe needs#
The honest answer depends on the recipe, because gpt-oss ships its experts in MXFP4 and only some tools can train on top of that. The 13.8 GB checkpoint becomes 41.8 GB of weights once it is unpacked to BF16 (our calculation: 20.91 billion parameters x 2 bytes).
| Recipe | What trains | GPU memory | Source |
|---|---|---|---|
| OpenAI cookbook: Transformers, PEFT and TRL (this guide) | LoRA on BF16 weights: attention layers and the experts of layers 7, 15 and 23 | Ran on one H100, which has 80 GB in its SXM and PCIe versions | OpenAI cookbook |
| Unsloth | QLoRA on a 4-bit base | 14 GB | Unsloth docs |
| Unsloth | 16-bit LoRA | 44 GB | Unsloth docs |
| Axolotl example | LoRA on the linear layers, 4,096-token packed rows, 8-bit AdamW, activations offloaded to CPU | About 44 GiB on one 48 GB GPU | Axolotl's gpt-oss README |
| Any library that upcasts to BF16, per Unsloth | Training in BF16 | At least 65 GB | Unsloth docs |
That rules a 48 GB card out for this recipe: the BF16 weights leave about 6 GB for everything else (our calculation: 48 - 41.8). On a 48 GB GPU, Unsloth's QLoRA is the route, and the Unsloth guide covers its workflow. Unsloth's gpt-oss notebook pins older Transformers and TRL releases than this recipe, so give it an environment of its own. Axolotl on a multi-GPU VM covers installing Axolotl, whose gpt-oss example is the table's other 48 GB figure. The LoRA, QLoRA and full fine-tuning VRAM guide compares these figures with other models.
Which GPU to rent#
The rule I follow here is to take the lowest-priced GPU with 80 GB or more, because the recipe is short and memory is what it needs. The single-GPU options with 80 GB or more on 2026-09-27:
| GPU | Memory | Pick it when |
|---|---|---|
| A100 SXM4 80GB | 80 GB | Price matters most. As much memory as an H100 SXM or PCIe |
| H100 PCIe | 80 GB | You want speed: 756 dense BF16 TFLOPS against the A100's 312 |
| RTX PRO 6000 Blackwell | 96 GB | You want 16 GB of headroom for longer sequences or a bigger batch |
| H200 NVL | 141 GB | You also plan to fine-tune or serve gpt-oss-120b |
Their live prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| A100 SXM4 80GB | 80 GB | $1.50/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 18:09 UTC
The harmony chat format#
gpt-oss was trained on OpenAI's harmony format, and fine-tuning data has to arrive in it. You do not write the special tokens yourself: each example is a list of messages, and the tokenizer's chat template renders them. In the dataset used here, every row has a messages list, and each assistant turn carries a thinking field for the reasoning next to content for the answer. This is what the template makes of a short conversation, rendered with the current tokenizer:
<|start|>system<|message|>You are ChatGPT, a large language model trained by OpenAI.
Knowledge cutoff: 2024-06
Current date: 2026-09-28
Reasoning: medium
# Valid channels: analysis, commentary, final. Channel must be included for every message.<|end|><|start|>developer<|message|># Instructions
reasoning language: German
<|end|><|start|>user<|message|>What is the capital of Australia?<|end|><|start|>assistant<|channel|>analysis<|message|>Die Frage ist einfach: Die Hauptstadt ist Canberra.<|end|><|start|>assistant<|channel|>final<|message|>The capital of Australia is Canberra.<|return|>
Each message ends with <|end|>, except the final answer, which ends with <|return|>. Four parts matter for training:
| Part | What it is |
|---|---|
system block | Written by the template: the model identity, the date and Reasoning: medium, the default reasoning effort |
developer block | Your system message goes here, as instructions. The dataset uses it to set the reasoning language |
analysis channel | The thinking field: the model's chain of thought |
final channel | The content field: the answer shown to the user |
OpenAI's cookbook trains on HuggingFaceH4/Multilingual-Thinking, 1,000 conversations under Apache-2.0, 200 each with the reasoning in English, French, German, Italian and Spanish, so the model learns to think in the language the developer message asks for. By our count with the gpt-oss chat template it is 1.14 million tokens, and 43 conversations run past 2,048 tokens and are cut there. To keep the model's reasoning, Unsloth's advice is to make at least 75% of your own examples reasoning examples like these.
Launch the GPU and install#
Launch an A100 SXM4 with the Bare Metal template, and connect as ubuntu (SSH docs). Your balance has to cover the first hour to launch (adding credits).
Check the GPU and driver with nvidia-smi, then create an environment with the releases this recipe is written for:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/gptoss-env --python 3.12
source ~/gptoss-env/bin/activate
uv pip install torch --torch-backend=auto
uv pip install "transformers==5.17.0" "peft==0.21.0" "trl==1.14.0" accelerate datasets
hf auth login
--torch-backend=auto makes uv pick the PyTorch build that matches the installed driver. Log in with a Hugging Face token that has write permission, for the upload at the end. The model is Apache-2.0 and not gated.
The training script#
The script is the cookbook's with small changes for the current releases: dtype replaces torch_dtype, warmup_steps=0.03 replaces warmup_ratio, which Transformers 5 removed, logs go to the console instead of Trackio, the adapter is saved locally for hf upload instead of pushed from the script, and it prints the training time and peak memory at the end. Save it as train.py:
# train.py: LoRA fine-tune of gpt-oss-20b, following OpenAI's cookbook recipe
import torch
from datasets import load_dataset
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM, AutoTokenizer, Mxfp4Config
from trl import SFTConfig, SFTTrainer
MODEL = "openai/gpt-oss-20b"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
dataset = load_dataset("HuggingFaceH4/Multilingual-Thinking", split="train")
model = AutoModelForCausalLM.from_pretrained(
MODEL,
dtype=torch.bfloat16,
quantization_config=Mxfp4Config(dequantize=True), # MXFP4 experts to BF16, so they can train
attn_implementation="eager",
use_cache=False,
device_map="auto",
)
peft_config = LoraConfig(
r=8,
lora_alpha=16,
target_modules="all-linear",
target_parameters=[ # the experts of three of the 24 layers
"7.mlp.experts.gate_up_proj", "7.mlp.experts.down_proj",
"15.mlp.experts.gate_up_proj", "15.mlp.experts.down_proj",
"23.mlp.experts.gate_up_proj", "23.mlp.experts.down_proj",
],
)
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()
args = SFTConfig(
output_dir="gpt-oss-20b-lora",
learning_rate=2e-4,
num_train_epochs=1,
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
gradient_checkpointing=True,
max_length=2048,
warmup_steps=0.03, # below 1, it is a fraction of all steps
lr_scheduler_type="cosine_with_min_lr",
lr_scheduler_kwargs={"min_lr_rate": 0.1},
logging_steps=1,
report_to="none",
)
trainer = SFTTrainer(model=model, args=args, train_dataset=dataset, processing_class=tokenizer)
torch.cuda.reset_peak_memory_stats()
result = trainer.train()
print(f"train_runtime_s = {result.metrics['train_runtime']:.0f}")
print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved() / 2**30:.1f}")
trainer.save_model("gpt-oss-20b-lora") # the LoRA adapter, not the base model
The settings that decide memory and results:
| Setting | Value | Why |
|---|---|---|
Mxfp4Config(dequantize=True) | on | Training needs BF16 weights. Transformers unpacks the experts while loading |
attn_implementation | eager | The cookbook's choice, and it runs on any GPU, the A100 included |
target_modules, target_parameters | all linear layers, plus three layers' experts | gpt-oss stores its experts as plain parameters, so PEFT targets them by name |
r, lora_alpha | 8, 16 | The cookbook's values. A higher rank trains more parameters and needs more memory |
per_device_train_batch_size, gradient_accumulation_steps | 4, 4 | 16 conversations per step, 63 steps for the epoch (our calculation: 1,000 / 16, rounded up) |
max_length | 2048 | Longer conversations are cut, which here affects 43 of 1,000 |
Run it and watch memory#
Run the script in the background, keep the log, and watch the GPU from a second terminal:
nohup python train.py > train.log 2>&1 &
tail -f train.log
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv -l 5
The first minutes go to downloading the 13.8 GB checkpoint and unpacking it to BF16. The loss is logged at every step. At the end, the script prints the training time and peak memory and saves the adapter.
The cookbook's 18 minutes on an H100 work out to about 990 tokens per second for this dataset (our calculation: 1,065,193 tokens / 1,080 seconds). The cookbook does not say which H100 it used. The cost guide turns figures like these into dollars.
Save the adapter off the VM#
Stopping a QuantaCloud instance terminates it and deletes its disk, so push the adapter before anything else. The folder holds the LoRA weights, the tokenizer files and a checkpoint-63 folder that the Trainer writes when training ends, with the optimizer state for resuming, but not the 20B base:
hf upload your-username/gpt-oss-20b-lora ./gpt-oss-20b-lora . --private
To keep a copy on your own machine as well, moving files to and from a GPU server covers scp, rsync and rclone.
Merge it and serve it with vLLM#
Serving engines want one set of weights, so fold the adapter into a BF16 copy of the base model. Loading with Mxfp4Config(dequantize=True) also drops the quantization config, so the saved folder is a plain BF16 checkpoint of about 41.8 GB. Save this as merge.py and run it with python merge.py:
# merge.py: fold the adapter into a BF16 copy of gpt-oss-20b for serving
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer, Mxfp4Config
base = AutoModelForCausalLM.from_pretrained(
"openai/gpt-oss-20b",
dtype=torch.bfloat16,
quantization_config=Mxfp4Config(dequantize=True),
device_map="auto",
)
model = PeftModel.from_pretrained(base, "gpt-oss-20b-lora").merge_and_unload()
model.save_pretrained("gpt-oss-20b-merged")
AutoTokenizer.from_pretrained("openai/gpt-oss-20b").save_pretrained("gpt-oss-20b-merged")
vLLM 0.30.0 supports gpt-oss and loads weights without a quantization config as plain BF16. Serve the merged folder with vLLM's Docker image, using the tag and API key setup from the vLLM Docker guide (v0.30.0 needs an R580 or newer driver, v0.30.0-cu129 covers older ones):
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/gpt-oss-20b-merged:/models/gpt-oss-20b-ft \
"vllm/vllm-openai:$VLLM_TAG" \
/models/gpt-oss-20b-ft \
--served-model-name gpt-oss-20b-ft \
--api-key "$VLLM_API_KEY"
On an 80 GB GPU, vLLM's default 92% budget leaves about 34.4 to 34.7 GiB after the weights, room for at most 1.5 million tokens of KV cache, or roughly 11 full 131,072-token conversations, before activations and CUDA graphs take their share (our calculation: 0.92 x the 81,559 to 81,920 MiB an 80 GB card reports, minus 41.8 GB, at 24,576 bytes per token for the 12 full-attention layers). Send the same kind of developer instruction you trained with:
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $VLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-oss-20b-ft", "messages": [{"role": "system", "content": "reasoning language: German"}, {"role": "user", "content": "What is the capital of Australia?"}], "max_tokens": 400}'
Chat Completions returns the reasoning separately from the final answer. The gpt-oss GPU requirements guide covers serving the original models, and it applies to the merged one. Push gpt-oss-20b-merged off the VM with hf upload too, or rebuild it later from the adapter.
What the session costs#
A session costs its length times the hourly price: the first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. At the A100 SXM4's price on 2026-09-27, $1.50 an hour, a 90-minute session is charged $1.50 at launch and $1.50 more at the one-hour mark, and $0.75 comes back for the unused half hour when you stop, so it costs $2.25 (our calculation: 2 x $1.50 - 0.5 x $1.50). Today's price is $1.50/GPU-hr. Pricing has the full rules.
FAQ#
Can I fine-tune gpt-oss-20b on a 48 GB GPU?
Yes, with a different recipe. Unsloth's QLoRA needs 14 GB and its 16-bit LoRA 44 GB, and Axolotl's example lists about 44 GiB. This guide's recipe unpacks the model to 41.8 GB of BF16 weights, which leaves too little of a 48 GB card, so it wants 80 GB.
Why not train the MXFP4 weights directly?
The MXFP4 kernels have no backward pass, according to Unsloth, and Transformers marks MXFP4 as not trainable and points to Mxfp4Config(dequantize=True). Every recipe here either unpacks to BF16 or, in Unsloth's case, converts the experts to 4-bit NF4 for QLoRA.
Should I train only on the assistant's answers?
That is TRL's assistant_only_loss=True, and TRL 1.14 swaps in a gpt-oss template that marks the assistant turns for it. The cookbook trains on whole conversations, as this script does. Try both on your own data.
What about gpt-oss-120b?
Unsloth lists QLoRA at 65 GB and 16-bit LoRA at 210 GB, so QLoRA fits one 80 GB GPU or the H200 NVL, and 16-bit LoRA needs several GPUs, such as 2x H200 NVL. Full fine-tuning of the 120b is a cluster job: we order and build GPU clusters to your spec and quote the configuration, lead time and terms in writing, send a capacity brief.
My rule for gpt-oss-20b: run the cookbook recipe on the lowest-priced GPU with 80 GB or more, which in the 2026-09-27 catalog was the A100 SXM4, push the adapter off the VM the moment training ends, and merge only when you need to serve it. If you only have a 48 GB card in mind, use Unsloth's QLoRA instead. The fine-tuning overview links the other recipes.
Launch an A100 SXM4 80GB