Your workload. Your GPU.

Make the model your own.

Fine-tune open LLMs on NVIDIA GPU VMs with LoRA, QLoRA or full fine-tuning: memory by model size, which GPU, and Unsloth, Axolotl, TRL, DeepSpeed or FSDP.

Fine-tuning an LLM on QuantaCloud means renting an NVIDIA GPU VM, running an open-source trainer on it, such as Unsloth, Axolotl or Hugging Face TRL, and copying the adapter or checkpoints off before you stop. QLoRA fits most jobs on one GPU: the published tables put a 7B to 8B model at 5 to 14 GB, a 32B model at 24 to 32 GB and a 70B model at 40 to 48 GB. There is no managed fine-tuning service or API. You run the training yourself, on one VM with 1, 2, 4 or 8 GPUs.

A 48 GB RTX A6000, enough for QLoRA up to the 32B class, runs from $0.48/GPU-hr, and the 141 GB H200 NVL from the console price. You pay from launch to stop, and the unused seconds of the current hour are refunded when you stop. Prices checked 6 Oct 2026, 00:22 UTC

Pick the method first#

The method sets the memory, and the memory sets the GPU.

Model sizeQLoRA (4-bit base)LoRA (16-bit base)Full fine-tuning, model states only
7B to 8B5 to 14 GB16 to 24 GB112 to 144 GB
13B to 14B8.5 to 20 GB28 to 40 GB208 to 252 GB
30B to 32B24 to 32 GB64 to 80 GB480 to 576 GB
70B40 to 48 GB160 to 164 GB1,120 to 1,260 GB

The QLoRA and LoRA ranges span the tables Unsloth, Axolotl and LLaMA-Factory publish, which are minimums or short-sequence estimates, so longer sequences and bigger batches need more. The full fine-tuning column is our calculation at 16 to 18 bytes per parameter for weights, gradients and Adam states, before activations. LoRA vs QLoRA vs full fine-tuning VRAM puts the three tables side by side and shows how sequence length moves them.

QLoRA keeps the base model frozen in 4-bit and trains a small adapter next to it, LoRA does the same on a 16-bit base, and full fine-tuning updates every weight. The rule I follow is to start with QLoRA and move up only for a reason. Unsloth's own advice is to test with LoRA or QLoRA first, because a task that fails there will almost certainly fail with full fine-tuning too, and a QLoRA run costs the least to find out.

Which GPU for which job#

Most runs start on a 48 GB card, and the table shows where each job moves up.

JobWhere I would run itWhy
QLoRA up to 32BOne 48 GB GPU: RTX A6000, RTX 6000 Ada, L40 or L40S24 to 32 GB at 32B leaves room for longer sequences
16-bit LoRA up to 14BOne 48 GB GPU28 to 40 GB at 14B
QLoRA at 70BOne A100 80GB, H100 PCIe or H200 NVLUnsloth reached 89,389 tokens of context on 80 GB, against 12,106 on 48 GB
16-bit LoRA at 32BOne RTX PRO 6000 (96 GB) or H200 NVL (141 GB)64 to 80 GB makes an 80 GB card tight
16-bit LoRA at 70B2x H200 NVL (282 GB) or 4x A100 80GB (320 GB)160 to 164 GB
Full fine-tuning, 7B to 14B2x H200 NVL for 8B, 4x A100 80GB or 4x RTX PRO 6000 for 14B, sharded112 to 252 GB of model states
Full fine-tuning at 70BA cluster built to order1.1 to 1.3 TB of model states, more than the largest on-demand VM

The 48 GB cards are the RTX A6000, RTX 6000 Ada, L40 and L40S, the 80 GB cards the A100 and H100 PCIe, and the 96 GB card the RTX PRO 6000 Blackwell. Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

Every on-demand instance is one VM with 1, 2, 4 or 8 GPUs, and there is no multi-node training on demand. On 2026-09-27 the largest was 8x A100 SXM4, with 640 GB of GPU memory and 800 GB of RAM. The offers API flags NVLink on the A100 SXM4 and H200 NVL offers, and sharded training moves weights and gradients between GPUs on every step, so check the links on your own VM with nvidia-smi topo -m before a run depends on them.

Launch 2x H200 NVL (282 GB) for 70B LoRA

The tools, and when to use each#

All five of these run on a QuantaCloud VM, and the choice matters less than the method.

ToolLatest release on 2026-09-28What it doesAcross several GPUs
Unsloth2026.9.12LoRA, QLoRA, full fine-tuning, DPO and GRPO, with its own memory-saving kernelsData parallel with torchrun or accelerate launch, and device_map = "balanced" to split one model
Axolotl0.19.0Full, LoRA and QLoRA fine-tuning plus DPO and GRPO, configured in one YAML fileDDP by default, DeepSpeed ZeRO 1 to 3, and FSDP2, which it recommends over FSDP1
Hugging Face TRL1.14.0Trainers for SFT, DPO, GRPO and reward models, called from your own Python codeThrough Accelerate, with ready-made configs for DDP, DeepSpeed ZeRO and FSDP
DeepSpeed0.19.7ZeRO sharding: optimizer states (stage 1), then gradients (stage 2), then weights (stage 3)Across all the GPUs of the VM
PyTorch FSDPPyTorch 2.14.0FSDP2 shards parameters, gradients and optimizer states across GPUsAcross all the GPUs of the VM

My default is Unsloth for anything that fits one GPU, Axolotl when a run needs several GPUs or a config file you can rerun, and TRL directly when the trainer belongs inside your own code. DeepSpeed and FSDP are what Axolotl and TRL use underneath to shard a model. ZeRO stage 3 and FSDP both split weights, gradients and optimizer states evenly: a full fine-tune of an 8B model holds 128 GB of model states at 16 bytes per parameter, which is 16 GB per GPU on eight GPUs before activations (our calculation: 8 x 16 = 128, and 128 / 8 = 16).

Give each tool its own environment. Their newest releases pin different library versions: Unsloth 2026.9.12 needs torch below 2.13 and Transformers up to 5.5.0, while Axolotl 0.19.0 pins Transformers 5.16.1 exactly, so one shared environment will not resolve. The Unsloth tutorial shows the pattern with uv, Axolotl on one 8-GPU VM covers the multi-GPU case, DeepSpeed ZeRO and DDP and FSDP cover the sharding, and fine-tuning gpt-oss-20b walks through one model end to end.

What a run costs#

A run costs its hours times the hourly price, and the hours come from your own seconds per step: hours = steps x seconds per step / 3,600, where steps per epoch on one GPU = examples / (batch size x accumulation steps). 3,000 examples at a batch of 1 with 4 accumulation steps make 750 steps per epoch (our calculation: 3,000 / 4).

At the prices on 2026-09-27, a 3-hour run comes to $1.44 on an RTX A6000 and $10.29 on an H200 NVL (our calculations: 3 x $0.48 and 3 x $3.43). Today's prices are $0.48/GPU-hr and the console price.

You pay for the time the VM exists, idle minutes included. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. If your balance cannot cover the next hour, the instance is terminated and its disk deleted with your checkpoints on it, so deposit enough for the whole run or turn on auto top-up (billing docs). What fine-tuning costs works through full examples, and low GPU utilization covers the fixes that shorten a run.

Keep your checkpoints when the VM ends#

The safe habit is to push checkpoints off the VM while the run is still going. Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots to fall back on. Save checkpoints along the way (save_steps in the trainer config) and upload the output folder on a schedule from a second terminal, in the trainer's environment, where the hf command comes with the Hugging Face libraries. Log in once with a write token, and the upload then pushes the folder to a private Hugging Face repo every 15 minutes until you press Ctrl+C:

Terminal
hf auth login
hf upload your-username/my-run ./outputs . --private --every=15

For a copy on your own machine, rsync over SSH transfers only what changed each time you run it, and moving files to and from a GPU server covers rsync, scp and rclone. On Bare Metal, start long runs with nohup or inside tmux, so a dropped SSH connection does not stop the training.

Templates for fine-tuning#

Two templates suit fine-tuning, and neither ships a trainer. PyTorch + Jupyter opens JupyterLab behind your QuantaCloud login, with a terminal for installs and a file browser for downloads. Bare Metal gives you Ubuntu 22.04 with the NVIDIA driver and Docker over SSH as ubuntu, which suits long runs and trainers that ship as Docker images, such as Axolotl. Templates compares all four, and the templates docs list what each one runs.

Launch an A100 80GB with Bare Metal for 70B QLoRA

When to reserve capacity instead#

A full fine-tune of a 70B model is a cluster job. Its weights, gradients and Adam states need 1.12 to 1.26 TB before activations, which is at least 15 GPUs with 80 GB each (our calculation: 70 billion x 16 to 18 bytes, and 1,260 GB is 1,173.5 GiB, 14.7 cards of the 79.65 to 80.00 GiB each reports), so it spans two 8-GPU servers and a fabric between them. A single 8-GPU B300 server holds the same job on one NVLink domain, with 2,160 GB of GPU memory (our calculation: 8 x 270 GB).

We order and build that hardware to your spec, from single 8-GPU servers to GPU clusters with InfiniBand, and quote the configuration, lead time and terms in writing before you commit. InfiniBand vs NVLink explains which link carries which traffic. Send a capacity brief

Questions before you fine-tune#

Is there a managed fine-tuning API?

No. QuantaCloud has no fine-tuning service, API or trainer template. You launch a GPU VM, install the trainer and run it, and the adapter or checkpoints are yours to copy off.

Can I train across several VMs?

Not on demand. Each on-demand instance is one VM with up to 8 GPUs, and there is no multi-node option, so multi-node training runs on reserved capacity built to order.

Rarely. With an adapter and plain data-parallel training, only the adapter's gradients cross between GPUs, and Axolotl puts the adapter at 0.1 to 1% of the parameters. Full sharding with FSDP or ZeRO stage 3 moves weights and gradients on every step, and that is where the link between GPUs starts to matter.

Can I fine-tune gpt-oss?

Yes. Unsloth lists QLoRA at 14 GB for gpt-oss-20b and 65 GB for gpt-oss-120b, so the 20b fits any 48 GB card and the 120b wants an 80 GB card or the H200 NVL. The gpt-oss GPU requirements guide sizes the GPU for serving the result, and self-hosted LLM inference compares the servers that run it.


My rule: start with QLoRA on the smallest GPU that holds the model with room to spare, which is a 48 GB card up to 32B and an 80 GB card or the H200 NVL at 70B, and read the peak memory the run prints before you size up. Move to 16-bit LoRA or full fine-tuning only when QLoRA falls short on your own evaluation, and push every checkpoint off the VM before you stop. The Unsloth tutorial is the first run I would do.

Launch a fine-tuning GPU

Keep building

Choose your next step.