Fine-tuning an LLM on QuantaCloud means renting an NVIDIA GPU VM, running an open-source trainer on it, such as Unsloth, Axolotl or Hugging Face TRL, and copying the adapter or checkpoints off before you stop. QLoRA fits most jobs on one GPU: the published tables put a 7B to 8B model at 5 to 14 GB, a 32B model at 24 to 32 GB and a 70B model at 40 to 48 GB. There is no managed fine-tuning service or API. You run the training yourself, on one VM with 1, 2, 4 or 8 GPUs.
A 48 GB RTX A6000, enough for QLoRA up to the 32B class, runs from $0.48/GPU-hr, and the 141 GB H200 NVL from the console price. You pay from launch to stop, and the unused seconds of the current hour are refunded when you stop. Prices checked 6 Oct 2026, 00:22 UTC
Pick the method first#
The method sets the memory, and the memory sets the GPU.
| Model size | QLoRA (4-bit base) | LoRA (16-bit base) | Full fine-tuning, model states only |
|---|---|---|---|
| 7B to 8B | 5 to 14 GB | 16 to 24 GB | 112 to 144 GB |
| 13B to 14B | 8.5 to 20 GB | 28 to 40 GB | 208 to 252 GB |
| 30B to 32B | 24 to 32 GB | 64 to 80 GB | 480 to 576 GB |
| 70B | 40 to 48 GB | 160 to 164 GB | 1,120 to 1,260 GB |
The QLoRA and LoRA ranges span the tables Unsloth, Axolotl and LLaMA-Factory publish, which are minimums or short-sequence estimates, so longer sequences and bigger batches need more. The full fine-tuning column is our calculation at 16 to 18 bytes per parameter for weights, gradients and Adam states, before activations. LoRA vs QLoRA vs full fine-tuning VRAM puts the three tables side by side and shows how sequence length moves them.
QLoRA keeps the base model frozen in 4-bit and trains a small adapter next to it, LoRA does the same on a 16-bit base, and full fine-tuning updates every weight. The rule I follow is to start with QLoRA and move up only for a reason. Unsloth's own advice is to test with LoRA or QLoRA first, because a task that fails there will almost certainly fail with full fine-tuning too, and a QLoRA run costs the least to find out.
Which GPU for which job#
Most runs start on a 48 GB card, and the table shows where each job moves up.
| Job | Where I would run it | Why |
|---|---|---|
| QLoRA up to 32B | One 48 GB GPU: RTX A6000, RTX 6000 Ada, L40 or L40S | 24 to 32 GB at 32B leaves room for longer sequences |
| 16-bit LoRA up to 14B | One 48 GB GPU | 28 to 40 GB at 14B |
| QLoRA at 70B | One A100 80GB, H100 PCIe or H200 NVL | Unsloth reached 89,389 tokens of context on 80 GB, against 12,106 on 48 GB |
| 16-bit LoRA at 32B | One RTX PRO 6000 (96 GB) or H200 NVL (141 GB) | 64 to 80 GB makes an 80 GB card tight |
| 16-bit LoRA at 70B | 2x H200 NVL (282 GB) or 4x A100 80GB (320 GB) | 160 to 164 GB |
| Full fine-tuning, 7B to 14B | 2x H200 NVL for 8B, 4x A100 80GB or 4x RTX PRO 6000 for 14B, sharded | 112 to 252 GB of model states |
| Full fine-tuning at 70B | A cluster built to order | 1.1 to 1.3 TB of model states, more than the largest on-demand VM |
The 48 GB cards are the RTX A6000, RTX 6000 Ada, L40 and L40S, the 80 GB cards the A100 and H100 PCIe, and the 96 GB card the RTX PRO 6000 Blackwell. Live prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Every on-demand instance is one VM with 1, 2, 4 or 8 GPUs, and there is no multi-node training on demand. On 2026-09-27 the largest was 8x A100 SXM4, with 640 GB of GPU memory and 800 GB of RAM. The offers API flags NVLink on the A100 SXM4 and H200 NVL offers, and sharded training moves weights and gradients between GPUs on every step, so check the links on your own VM with nvidia-smi topo -m before a run depends on them.
The tools, and when to use each#
All five of these run on a QuantaCloud VM, and the choice matters less than the method.
| Tool | Latest release on 2026-09-28 | What it does | Across several GPUs |
|---|---|---|---|
| Unsloth | 2026.9.12 | LoRA, QLoRA, full fine-tuning, DPO and GRPO, with its own memory-saving kernels | Data parallel with torchrun or accelerate launch, and device_map = "balanced" to split one model |
| Axolotl | 0.19.0 | Full, LoRA and QLoRA fine-tuning plus DPO and GRPO, configured in one YAML file | DDP by default, DeepSpeed ZeRO 1 to 3, and FSDP2, which it recommends over FSDP1 |
| Hugging Face TRL | 1.14.0 | Trainers for SFT, DPO, GRPO and reward models, called from your own Python code | Through Accelerate, with ready-made configs for DDP, DeepSpeed ZeRO and FSDP |
| DeepSpeed | 0.19.7 | ZeRO sharding: optimizer states (stage 1), then gradients (stage 2), then weights (stage 3) | Across all the GPUs of the VM |
| PyTorch FSDP | PyTorch 2.14.0 | FSDP2 shards parameters, gradients and optimizer states across GPUs | Across all the GPUs of the VM |
My default is Unsloth for anything that fits one GPU, Axolotl when a run needs several GPUs or a config file you can rerun, and TRL directly when the trainer belongs inside your own code. DeepSpeed and FSDP are what Axolotl and TRL use underneath to shard a model. ZeRO stage 3 and FSDP both split weights, gradients and optimizer states evenly: a full fine-tune of an 8B model holds 128 GB of model states at 16 bytes per parameter, which is 16 GB per GPU on eight GPUs before activations (our calculation: 8 x 16 = 128, and 128 / 8 = 16).
Give each tool its own environment. Their newest releases pin different library versions: Unsloth 2026.9.12 needs torch below 2.13 and Transformers up to 5.5.0, while Axolotl 0.19.0 pins Transformers 5.16.1 exactly, so one shared environment will not resolve. The Unsloth tutorial shows the pattern with uv, Axolotl on one 8-GPU VM covers the multi-GPU case, DeepSpeed ZeRO and DDP and FSDP cover the sharding, and fine-tuning gpt-oss-20b walks through one model end to end.
What a run costs#
A run costs its hours times the hourly price, and the hours come from your own seconds per step: hours = steps x seconds per step / 3,600, where steps per epoch on one GPU = examples / (batch size x accumulation steps). 3,000 examples at a batch of 1 with 4 accumulation steps make 750 steps per epoch (our calculation: 3,000 / 4).
At the prices on 2026-09-27, a 3-hour run comes to $1.44 on an RTX A6000 and $10.29 on an H200 NVL (our calculations: 3 x $0.48 and 3 x $3.43). Today's prices are $0.48/GPU-hr and the console price.
You pay for the time the VM exists, idle minutes included. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. If your balance cannot cover the next hour, the instance is terminated and its disk deleted with your checkpoints on it, so deposit enough for the whole run or turn on auto top-up (billing docs). What fine-tuning costs works through full examples, and low GPU utilization covers the fixes that shorten a run.
Keep your checkpoints when the VM ends#
The safe habit is to push checkpoints off the VM while the run is still going. Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots to fall back on. Save checkpoints along the way (save_steps in the trainer config) and upload the output folder on a schedule from a second terminal, in the trainer's environment, where the hf command comes with the Hugging Face libraries. Log in once with a write token, and the upload then pushes the folder to a private Hugging Face repo every 15 minutes until you press Ctrl+C:
hf auth login
hf upload your-username/my-run ./outputs . --private --every=15
For a copy on your own machine, rsync over SSH transfers only what changed each time you run it, and moving files to and from a GPU server covers rsync, scp and rclone. On Bare Metal, start long runs with nohup or inside tmux, so a dropped SSH connection does not stop the training.
Templates for fine-tuning#
Two templates suit fine-tuning, and neither ships a trainer. PyTorch + Jupyter opens JupyterLab behind your QuantaCloud login, with a terminal for installs and a file browser for downloads. Bare Metal gives you Ubuntu 22.04 with the NVIDIA driver and Docker over SSH as ubuntu, which suits long runs and trainers that ship as Docker images, such as Axolotl. Templates compares all four, and the templates docs list what each one runs.
When to reserve capacity instead#
A full fine-tune of a 70B model is a cluster job. Its weights, gradients and Adam states need 1.12 to 1.26 TB before activations, which is at least 15 GPUs with 80 GB each (our calculation: 70 billion x 16 to 18 bytes, and 1,260 GB is 1,173.5 GiB, 14.7 cards of the 79.65 to 80.00 GiB each reports), so it spans two 8-GPU servers and a fabric between them. A single 8-GPU B300 server holds the same job on one NVLink domain, with 2,160 GB of GPU memory (our calculation: 8 x 270 GB).
We order and build that hardware to your spec, from single 8-GPU servers to GPU clusters with InfiniBand, and quote the configuration, lead time and terms in writing before you commit. InfiniBand vs NVLink explains which link carries which traffic. Send a capacity brief
Questions before you fine-tune#
Is there a managed fine-tuning API?
No. QuantaCloud has no fine-tuning service, API or trainer template. You launch a GPU VM, install the trainer and run it, and the adapter or checkpoints are yours to copy off.
Can I train across several VMs?
Not on demand. Each on-demand instance is one VM with up to 8 GPUs, and there is no multi-node option, so multi-node training runs on reserved capacity built to order.
Do I need NVLink for LoRA?
Rarely. With an adapter and plain data-parallel training, only the adapter's gradients cross between GPUs, and Axolotl puts the adapter at 0.1 to 1% of the parameters. Full sharding with FSDP or ZeRO stage 3 moves weights and gradients on every step, and that is where the link between GPUs starts to matter.
Can I fine-tune gpt-oss?
Yes. Unsloth lists QLoRA at 14 GB for gpt-oss-20b and 65 GB for gpt-oss-120b, so the 20b fits any 48 GB card and the 120b wants an 80 GB card or the H200 NVL. The gpt-oss GPU requirements guide sizes the GPU for serving the result, and self-hosted LLM inference compares the servers that run it.
My rule: start with QLoRA on the smallest GPU that holds the model with room to spare, which is a 48 GB card up to 32B and an 80 GB card or the H200 NVL at 70B, and read the peak memory the run prints before you size up. Move to 16-bit LoRA or full fine-tuning only when QLoRA falls short on your own evaluation, and push every checkpoint off the VM before you stop. The Unsloth tutorial is the first run I would do.
Launch a fine-tuning GPU