GPU guide

How much does it cost to fine-tune an LLM on rented GPUs? Worked examples

Estimate a fine-tune's cost: count the tokens, get tokens per second from a published or measured run, turn that into hours, and price them at live rates.

Faiz Ahmed10 min read

A fine-tune costs its hours on the GPU times the hourly price, and the hours come from three numbers you can pin down before you rent anything: the tokens in your dataset, the number of epochs, and the tokens per second your setup trains. Hours = tokens x epochs / (tokens per second x 3,600). The gpt-oss-20b LoRA run in OpenAI's cookbook shows how it works: its dataset is about 1.07 million tokens by our count, it trained in about 18 minutes on one H100, roughly 990 tokens per second, and if the H100 PCIe matched that speed, 18 minutes at its $2.59 an hour on 2026-09-27 would be $0.78 of training (our calculation: 0.3 x $2.59). Today the H100 PCIe is the console price. The session around the training, with installs, downloads and uploads, is billed too, so the bill always comes out higher than the training alone.

The estimate in five steps#

The formula needs no GPU until step 3, and even there a short test run is enough:

  1. Count the tokens in one epoch by running your dataset through the model's tokenizer and chat template.
  2. Choose the number of epochs.
  3. Get a tokens-per-second figure for your model, method and GPU: from a published run, or from a 50-step test on the GPU you plan to rent.
  4. Divide: hours = tokens x epochs / (tokens per second x 3,600). Add the setup time around the training.
  5. Multiply the hours by the VM's price per hour, which is the per-GPU price times the number of GPUs.

For step-based tools the same arithmetic works with steps: hours = steps x seconds per step / 3,600, with steps per epoch = examples / (batch size x accumulation steps x GPUs). The LoRA and QLoRA VRAM guide uses that form.

Step 1: count the tokens#

The core problem with estimating from example counts is that examples vary enormously in length: from 151 to 32,962 tokens in Multilingual-Thinking, by our count. Count tokens instead, with the model you will train, because the chat template adds tokens of its own. This script runs on any laptop, with no GPU:

Python
# count_tokens.py: tokens per epoch for a chat dataset under a model's chat template
import sys

from datasets import load_dataset
from transformers import AutoTokenizer

model, dataset, max_len = sys.argv[1], sys.argv[2], int(sys.argv[3])
tok = AutoTokenizer.from_pretrained(model)
rows = load_dataset(dataset, split="train")["messages"]
lengths = [len(tok(tok.apply_chat_template(m, tokenize=False), add_special_tokens=False)["input_ids"])
           for m in rows]
kept = sum(min(n, max_len) for n in lengths)
print(f"{len(lengths)} examples, {sum(lengths):,} tokens, {kept:,} after cutting each at {max_len}")

Run it as python count_tokens.py Qwen/Qwen3-8B trl-lib/Capybara 2048, after pip install transformers datasets. It expects a messages column of role and content turns. The last number matters because training cuts every example at the maximum sequence length. Our counts for the datasets used in QuantaCloud's fine-tuning guides:

DatasetChat templateConversationsTokensTokens after the cut
trl-lib/CapybaraQwen315,80615,060,54314,432,951 at 2,048
HuggingFaceH4/Multilingual-Thinkinggpt-oss1,0001,139,6561,065,193 at 2,048
mlabonne/FineTome-100k, first 20%Qwen320,00011,122,03811,079,495 at 4,096

FineTome-100k stores each turn as from and value, so we mapped those to role and content before counting. Some trainers drop over-long examples instead of cutting them: Axolotl does by default, which leaves 10,829,639 tokens of that FineTome slice, the figure the Axolotl guide uses.

Step 2: find tokens per second#

Tokens per second is the number that moves the estimate most, and it depends on the model size, LoRA or full fine-tuning, sequence length, batch size, the framework and the GPU. There are three ways to get it, from rough to reliable.

A published run

A published run with a known dataset gives a real figure. OpenAI's cookbook trains gpt-oss-20b with LoRA on Multilingual-Thinking and reports about 18 minutes on an H100, which is about 990 tokens per second (our calculation: 1,065,193 tokens / 1,080 seconds). It does not say which H100, and QuantaCloud's is the PCIe version, so I treat figures like this as a starting point, not a promise. The gpt-oss-20b fine-tuning guide runs the same recipe.

The compute ceiling

The compute ceiling is the speed no run can beat. Training costs about 6 floating-point operations per parameter per token, according to Kaplan and colleagues' scaling-law paper: 2 for the forward pass and 4 for the backward. Divide the GPU's dense BF16 throughput by that:

Full fine-tune of Qwen3-8B (8.19B parameters)Dense BF16 TFLOPS per GPUCeiling per GPUCeiling on 8 GPUs
A100 80GB312About 6,349 tokens/sAbout 50,789 tokens/s
H100 PCIe756About 15,383 tokens/sAbout 123,066 tokens/s

These are our calculations: 312 or 756 trillion operations per second divided by 6 x 8.19 billion, counting all of the model's parameters. No real run reaches the ceiling, because memory traffic, communication and recomputed activations all take time. Its use is as a floor on hours: an estimate that implies a faster run is wrong. LoRA does less work per token, because it skips the weight gradients for the frozen model, and the LoRA paper measured 25% faster training than full fine-tuning on GPT-3 175B.

Your own test run

A 50-step test on the GPU you plan to rent is the figure I would budget on. The scripts in the DDP and FSDP guide and the DeepSpeed ZeRO guide print tokens_per_s at the end, and the Hugging Face Trainer under TRL and Axolotl reports train_runtime and steps per second when it finishes. A test like that costs minutes of GPU time, because unused seconds are refunded when you stop.

Step 3: hours and dollars, worked examples#

These are our calculations at the prices on 2026-09-27. The first uses a published run and the second the compute ceiling, so treat both as floors: the session around the training adds time, and no run reaches the ceiling.

JobTokensTokens per secondHoursVM price per hour on 2026-09-27Cost
gpt-oss-20b LoRA, 1 epoch, one H100 PCIe1,065,193About 990, published run0.3$2.59$0.78 or more
Qwen3-8B full fine-tune, 3 epochs of Capybara, 8x A100 SXM443,298,853At most 50,789, compute ceiling0.237 or more$11.92$2.82 or more

Live prices per GPU-hour for the GPUs in these examples:

GPUMemoryFromAvailable now
H100 PCIe-Not listedNo
A100 SXM4 80GB80 GB$1.50/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX A600048 GB$0.48/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 16:18 UTC

How billing turns hours into charges#

You pay for the seconds between launch and stop, prepaid one hour at a time. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. A run of 2 hours 20 minutes on 8x A100 SXM4 at $11.92 an hour on 2026-09-27, which is $1.49 per GPU-hour (today: $1.50/GPU-hr), goes like this (our calculation):

Time after launchWhat happensBalance change
0:00Launch: hour 1 charged-$11.92
1:00Hour 2 charged-$11.92
2:00Hour 3 charged-$11.92
2:20Stop: 40 unused minutes refunded (40/60 x $11.92)+$7.95
Net2.333 hours x $11.92$27.81

The part that needs planning is the balance, not the rate. If your balance cannot cover the next hour, the instance is terminated and its disk deleted, checkpoints and all. For a long run I fund the whole estimate plus one hour first, or turn on auto top-up. Pricing has every rule, and the billing docs show the Billing page itself.

What the formula leaves out#

The formula counts training time only, and the VM bills for every minute around it as well:

ItemWhy it costs time
BootMost single-GPU VMs are running in about 3 minutes (median), and 8-GPU VMs in about 10. The clock starts at Deploy and those minutes count once the VM is running
Installs and downloadsA fresh VM has no Python environment and no model weights. Every launch downloads the model again, because stopping deletes the disk
Evaluation and checkpointsSaving and uploading an 8B checkpoint with optimizer state moves 98 GB or more
Failed first attemptsOut-of-memory errors and wrong settings cost least when a short test run finds them
Idle GPUsA VM left running after the job finishes bills until you stop it

A small run shows the scale. The Unsloth guide prices a 45-minute session on an RTX A6000 at $0.48 an hour on 2026-09-27: $0.48 is charged at launch and the unused 15 minutes come back at stop, so it costs $0.36 (our calculation: 0.75 x $0.48). The RTX A6000 is $0.48/GPU-hr today. To stop paying for idle GPUs, the auto-stop guide stops an instance through the API once its GPUs go idle, and low GPU utilization covers the slow runs that inflate the hours.

When to ask for a reserved build#

Hourly billing suits runs that start, finish and stop. The rule on the pricing page applies here too: if you would keep the same GPUs busy most hours of most weeks, a reserved quote is worth setting against the arithmetic above. The same goes for a job that needs more GPUs than one VM holds, such as a full fine-tune of a 70B model, whose model states alone are 1.1 to 1.3 TB. We order and build that hardware to your spec, from single servers to GPU clusters, and quote the configuration, lead time and terms in writing: send a capacity brief.

FAQ#

How much does it cost to fine-tune a 7B or 8B model?

For QLoRA or LoRA, the GPU is one 48 GB card, since the published needs are 5 to 14 GB for QLoRA and 16 to 24 GB for LoRA. On 2026-09-27 the RTX A6000 was $0.48 an hour (now $0.48/GPU-hr), so the cost is your hours at that rate. The hours come from your tokens and the step time of a short test run. A full fine-tune needs 112 to 144 GB of model states before activations, which means several GPUs sharded with FSDP or ZeRO.

Does a 10-minute test cost a whole hour?

No. The first hour is charged at launch, and the unused seconds come back when you stop, so a 10-minute test costs 10 minutes of the hourly rate.

How much does it cost to train an LLM from scratch?

That is pre-training, a different scale. Meta's model card lists 1.46 million H100 GPU-hours to train Llama 3.1 8B, a model pre-trained on about 15 trillion tokens. Fine-tuning the same size of model on a few million tokens uses the same formula, and takes GPU-hours rather than millions of them.

Is LoRA cheaper than full fine-tuning?

Usually, mostly through memory. The LoRA paper reports 3 times less GPU memory and 25% faster training than full fine-tuning on GPT-3 175B. Less memory means fewer or smaller GPUs, and on QuantaCloud that means a lower price per hour.


My rule for budgeting a fine-tune: count the tokens, run 50 steps on the GPU you plan to rent, multiply the measured time out to the full run, and add the setup around it. Fund the whole estimate plus one hour, push checkpoints off the VM as you go, and stop the instance the moment the job ends. The fine-tuning overview links the recipes behind every figure on this page.

Launch an RTX A6000 for a test run

Keep building

Choose your next step.