GPU guide

How much VRAM do you need to fine-tune an LLM?

How much GPU memory QLoRA, LoRA and full fine-tuning need for 7B to 70B models, why published tables disagree, and which GPU setups fit.

Faiz Ahmed16 min read

QLoRA fits most fine-tuning jobs on one GPU: the published tables put a 7B to 8B model at 5 to 14 GB, a 32B model at 24 to 32 GB and a 70B model at 40 to 48 GB. LoRA on a 16-bit base needs 16 to 24 GB at 8B and 160 to 164 GB at 70B. Full fine-tuning with mixed-precision Adam needs 16 to 18 bytes per parameter before activations, which puts a 70B model at 1.1 to 1.3 TB (our calculation: 70 billion x 16 to 18 bytes), more than any single on-demand QuantaCloud VM holds. The tables from Unsloth, Axolotl and LLaMA-Factory disagree because they assume different sequence lengths, batch sizes and optimizers, so every figure below carries its conditions and the QuantaCloud GPU instance it fits.

The short answer by model size#

The table I plan with takes the range across the three published tables for each method. The full fine-tuning column is our calculation from bytes per parameter (explained in the next section), because the published full fine-tuning figures mix two precision regimes.

Model sizeQLoRA (4-bit base)LoRA (16-bit base)Full fine-tuning, model states only
7B to 8B5 to 14 GB16 to 24 GB112 to 144 GB (our calculation)
13B to 14B8.5 to 20 GB28 to 40 GB208 to 252 GB (our calculation)
30B to 32B24 to 32 GB64 to 80 GB480 to 576 GB (our calculation)
70B40 to 48 GB160 to 164 GB1,120 to 1,260 GB (our calculation)

The QLoRA and LoRA ranges span the three tables compared further down. The full fine-tuning column is the parameter count x 16 bytes (low end) to x 18 bytes (high end).

Here is where each case runs on QuantaCloud. The mapping is ours, from the catalog on 2026-09-27, and it picks the configuration that holds the published figure with room to spare, because longer sequences and bigger batches eat that room quickly.

Model sizeQLoRALoRA (16-bit)Full fine-tuning
7B to 8BAny 48 GB GPUAny 48 GB GPU2x RTX PRO 6000 (192 GB) or 2x H200 NVL (282 GB), sharded
13B to 14BAny 48 GB GPUAny 48 GB GPU4x A100 80GB (320 GB) or 4x RTX PRO 6000 (384 GB), sharded
30B to 32BAny 48 GB GPURTX PRO 6000 (96 GB) or H200 NVL (141 GB), 80 GB is tight8x A100 SXM4 80GB (640 GB), sharded, tight at 32B
70BA100 80GB, H100 PCIe or H200 NVL, 48 GB only for short sequences2x H200 NVL (282 GB) or 4x A100 80GB (320 GB)Not on one on-demand VM: a cluster built to order

The 48 GB GPUs are the RTX A6000, RTX 6000 Ada, L40 and L40S. The 80 GB GPUs are the A100 80GB and the H100 PCIe, the 96 GB GPU is the RTX PRO 6000 Blackwell, and the 141 GB GPU is the H200 NVL. Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes
H100 PCIe-Not listedNo
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 18:28 UTC

Launch an RTX A6000 (48 GB) for QLoRA up to 32B Launch an H200 NVL (141 GB) for 70B QLoRA Launch 2x H200 NVL (282 GB) for 70B LoRA

Why fine-tuning needs more memory than inference#

The weights are the small part of a training job. Inference holds the weights and a KV cache. Training also holds a gradient for every trainable weight, the optimizer's state for each of those, and the activations the forward pass saves for the backward pass. Two accountings of full fine-tuning with Adam are in common use, plus a leaner regime:

AccountingWeightsGradientsOptimizer statesTotal per parameter
ZeRO paper, mixed precision Adam2 bytes (fp16)2 bytes (fp16)12 bytes (fp32 copy, momentum, variance)16 bytes
Hugging Face Transformers, mixed precision AdamW6 bytes (fp16 plus fp32 copy)4 bytes (fp32)8 bytes (momentum, variance)18 bytes
Pure bf16 with AdamW, no fp32 copy (LLaMA-Factory, Axolotl)not itemisednot itemisednot itemisedabout 8 bytes

Activations come on top and scale with batch size and sequence length. The Hugging Face memory guide gives the example of a 4B model trained in mixed precision at batch size 16, which needs roughly 85 GB. That is why a model that serves comfortably on a GPU can run out of memory the moment you train it.

LoRA and QLoRA change the arithmetic by freezing the base model. Gradients and optimizer states then exist only for the adapter, which Axolotl puts at about 0.1 to 1% of the parameters. What remains is the frozen base: 2 bytes per parameter in 16-bit, and in 4-bit NF4 with double quantization, 4 bits plus 0.127 bits of quantization constants per parameter, per the QLoRA paper. For a 70B model that is at least 140 GB for LoRA and 36.1 GB for QLoRA (our calculation: 70 billion x 2 bytes, and 70 billion x 4.127 bits / 8). Real checkpoints come in higher: Llama 3.3 70B is 141.1 GB in BF16, and 40.8 GB in the 4-bit version Unsloth loads, because 4-bit checkpoints keep the embeddings, the output layer and, in Unsloth's dynamic versions, some other layers in 16-bit (Hugging Face API, 2026-09-28). The published 160 to 164 GB and 40 to 48 GB are those frozen weights plus the adapter, its optimizer and the activations at short sequence lengths.

LoRA and QLoRA in one paragraph each#

LoRA freezes the pretrained weights and trains pairs of small low-rank matrices next to them. Its paper reports, for GPT-3 175B, 10,000 times fewer trainable parameters and 3 times less GPU memory than full fine-tuning with Adam. The adapter you save is small, and you either load it on top of the base model at serving time or merge it into the weights.

QLoRA keeps that frozen base in 4-bit NormalFloat (NF4), quantizes the quantization constants a second time, and uses paged optimizers to absorb memory spikes. The QLoRA paper's headline is that it can "finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance". The price is speed: Axolotl lists QLoRA as slower than LoRA because of the dequantization overhead.

The published tables, side by side#

The rule I follow is to plan on the Axolotl column, because it states its conditions, and to treat the Unsloth column as a floor you reach only with Unsloth's kernels at short sequence lengths.

Model sizeMethodUnslothAxolotlLLaMA-Factory
7B to 8BQLoRA5 GB (7B), 6 GB (8B)10 to 14 GB6 GB (7B)
7B to 8BLoRA19 GB (7B), 22 GB (8B)16 to 24 GB16 GB (7B)
7B to 8BFullnot listed60 to 80 GB120 GB (18x rule), 60 GB (pure bf16)
13B to 14BQLoRA8.5 GB (14B)16 to 20 GB12 GB (14B)
13B to 14BLoRA33 GB (14B)28 to 40 GB32 GB (14B)
13B to 14BFullnot listed120+ GB240 GB (18x rule), 120 GB (pure bf16)
30B to 32BQLoRA26 GB (32B)24 to 32 GB24 GB (30B)
30B to 32BLoRA76 GB (32B)64 to 80 GB64 GB (30B)
30B to 32BFullnot listed2x to 4x 80 GB600 GB (18x rule), 300 GB (pure bf16)
70BQLoRA41 GB40 to 48 GB48 GB
70BLoRA164 GB2x 80 GB160 GB
70BFullnot listed4x to 8x 80 GB1,200 GB (18x rule), 600 GB (pure bf16)

Sources: the Unsloth requirements page, Axolotl's method guide and the LLaMA-Factory README, all read on 2026-09-28.

The conditions explain the spread. Unsloth calls its numbers "the absolute minimum" and warns that some models need more. Axolotl's figures "assume a short context length (512-2048 tokens) and micro_batch_size of 1-2". LLaMA-Factory labels its table "estimated" and gives no sequence length, batch size or GPU count. Three things account for most of the gap. Sequence length and batch size drive the activations, and each table assumes different ones. Kernels matter too: in Unsloth's own benchmark an 8B QLoRA run on an 80 GB GPU reached 342,733 tokens of context, against 28,454 for Hugging Face with FlashAttention 2. And the full fine-tuning rows assume different regimes: LLaMA-Factory's 18x row keeps fp32 master weights, while its pure bf16 row and Axolotl's "roughly 4x model size in bf16 with AdamW" do not. Axolotl's Llama 3 8B example even runs a full fine-tune on a single 48 GB GPU, with an 8-bit paged optimizer and gradient checkpointing.

Context length changes everything#

Sequence length is the variable that turns a comfortable fit into an out-of-memory error. Unsloth measured the longest sequence a QLoRA run (rank 32, batch size 1, all linear layers) could train at each memory size, published on its benchmarks page:

GPU memoryLlama 3.1 8B, UnslothLlama 3.1 8B, Hugging Face + FA2Llama 3.3 70B, UnslothLlama 3.3 70B, Hugging Face + FA2
24 GB78,475 tokens5,789 tokensnot testednot tested
48 GB191,728 tokens15,502 tokens12,106 tokensout of memory
80 GB342,733 tokens28,454 tokens89,389 tokens6,916 tokens

The 48 GB rows are what the RTX A6000, RTX 6000 Ada, L40 and L40S give you, and the 80 GB rows are the A100 80GB and H100 PCIe. Unsloth did not publish rows for 96 or 141 GB. For a 70B model the step from 48 to 80 GB takes the maximum from 12,106 to 89,389 tokens, which is why I treat 48 GB as a short-context option for 70B QLoRA and 80 GB as the default. When a run does run out of memory, the CUDA out of memory guide orders the fixes by what they cost.

Multi-GPU: when one card is not enough#

One card is enough for every QLoRA job in the tables above. A second GPU becomes necessary for 16-bit LoRA at 70B and for full fine-tuning beyond small models. The ZeRO paper gives the memory for model states per GPU when N GPUs share a model of P parameters:

ShardingModel states per GPU
Plain data parallel, no sharding16 x P bytes
ZeRO stage 1 (optimizer states sharded)4 x P + 12 x P / N bytes
ZeRO stage 2 (gradients sharded too)2 x P + 14 x P / N bytes
ZeRO stage 3 (weights sharded too)16 x P / N bytes

A full fine-tune of an 8B model at 16 bytes per parameter holds 128 GB of model states, so ZeRO stage 3 on 8 GPUs puts 16 GB on each GPU before activations. At 70B the same split is 140 GB per GPU: nearly all of an H200 NVL's 141 GB before a single activation, and more than any other GPU in the catalog holds. It takes 16 GPUs to bring it to 70 GB each (our calculation: 16 x 70 / 16). That is a cluster, not a VM.

For LoRA and QLoRA at 70B, two recipes cover the multi-GPU case. Unsloth splits a model that does not fit one GPU with device_map = "balanced". Hugging Face PEFT measured QLoRA with FSDP on Llama 2 70B at sequence length 2,048 on two GPUs: 35.6 GB per GPU without CPU offload, and 19.6 GB per GPU with offload, which used about 107 GB of host RAM. At those settings the first figure fits a 2x 48 GB VM without any offload.

If you do plan on offload, check host RAM before GPU memory. On 2026-09-27 the 2x RTX A6000 configurations came with 48 to 128 GB of RAM, the 2x L40S with 144 GB and the 2x H200 NVL with 360 GB. The live RTX A6000 configurations:

GPUsvCPURAMDiskRegionPer hourPer GPU-hourLaunch
1x RTX A6000648 GB256 GBus-midwest-2$0.48$0.48Launch
1x RTX A6000648 GB256 GBus-midwest-1$0.48$0.48Launch
1x RTX A6000624 GB256 GBus-midwest-1$0.55$0.55Launch
1x RTX A6000624 GB256 GBus-midwest-2$0.55$0.55Launch
1x RTX A60001264 GB256 GBus-midwest-1$0.57$0.57Launch
1x RTX A60001264 GB256 GBus-midwest-2$0.57$0.57Launch
1x RTX A60001264 GB256 GBus-midwest-3$0.57$0.57Launch
2x RTX A60001496 GB512 GBus-midwest-1$0.96$0.48Launch
2x RTX A60001496 GB512 GBus-midwest-2$0.96$0.48Launch
2x RTX A60001448 GB512 GBus-midwest-2$1.08$0.54Launch
2x RTX A60001448 GB512 GBus-midwest-1$1.08$0.54Launch
2x RTX A600030128 GB512 GBus-midwest-1$1.13$0.56Launch
2x RTX A600030128 GB512 GBus-midwest-2$1.13$0.56Launch
4x RTX A600030192 GB1,024 GBus-midwest-1$1.92$0.48Launch
4x RTX A600030192 GB1,024 GBus-midwest-2$1.92$0.48Launch
4x RTX A60003096 GB1,024 GBus-midwest-1$2.16$0.54Launch

Worked examples on QuantaCloud#

Calculated: bigger runs

These are our calculations from the published figures above. For the frozen base I took a named model of each size and the checkpoint Unsloth loads for it, as listed on the Hugging Face API on 2026-09-28. None of these has been run on QuantaCloud.

RunFrozen basePublished needGPU and what is left over (our calculation)
32B QLoRA on 1x L40S19.2 GB (Qwen3-32B, 4-bit)24 to 32 GB48 GB leaves 16 to 24 GB for longer sequences or a bigger batch
70B QLoRA on 1x H200 NVL40.8 GB (Llama 3.3 70B, 4-bit)40 to 48 GB141 GB leaves 93 to 101 GB for long context
70B LoRA on 2x H200 NVL141.1 GB (Llama 3.3 70B, BF16)160 to 164 GB282 GB leaves 118 to 122 GB, with the model split across both GPUs
70B full fine-tuning1,120 to 1,260 GB of model states (16 to 18 bytes per parameter)1,200 GB (LLaMA-Factory), 4x to 8x 80 GB (Axolotl)More than the largest on-demand VM on 2026-09-27 (8x A100 SXM4, 640 GB)

A 70B full fine-tune is a cluster job. We order and build that hardware to your spec, from single servers to GPU clusters with InfiniBand, and quote the configuration, lead time and terms in writing: Send a capacity brief.

What a run costs#

A run costs its hours times the hourly price, and the hours come from your measured seconds per step. The formula I use is hours = steps x seconds per step / 3,600, where steps per epoch = examples / (batch size x accumulation steps). The Unsloth tutorial's 3,000 examples at a batch of 1 with 4 accumulation steps are 750 steps per epoch (our calculation: 3,000 / 4). Estimating fine-tuning cost works through more examples.

Multiply the hours by the live price: $0.48/GPU-hr for the RTX A6000 and the console price for the H200 NVL. At the prices on 2026-09-27, a 3-hour run cost $1.44 on an RTX A6000 and $10.29 on an H200 NVL (our calculations: 3 x $0.48 and 3 x $3.43). The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. Pricing has the full billing rules.

Two rules protect the result of a run. Stopping an instance terminates it and deletes its disk, so push checkpoints and the adapter to the Hugging Face Hub or your own object storage, or copy them off with scp, before you stop. And if your balance cannot cover the next hour, the instance is terminated and its disk deleted with it, so turn on auto top-up or deposit enough for the whole run first (billing docs).

Which method should you use#

The rule I follow is to start with QLoRA and move up only for a reason. It fits the most models on the fewest GPUs, and it is the quickest way to learn whether fine-tuning helps at all: Unsloth's advice is to test with LoRA or QLoRA first, because a task that fails there will almost certainly fail with full fine-tuning too.

Move to LoRA on a 16-bit base when the base fits in 16-bit with room to spare and you will serve the model in 16-bit. At 8B that is any 48 GB GPU, and at 32B it is the RTX PRO 6000 or the H200 NVL. Move to full fine-tuning when you are changing the model broadly, for example with continued pretraining, and you have measured that LoRA falls short. Up to 14B that runs on a multi-GPU VM, 32B only just fits the 8x A100 SXM4, and at 70B it is a cluster we build to order.

Common questions#

Can I fine-tune a 70B model on one GPU?

Yes, with QLoRA. The published need is 40 to 48 GB, so an 80 GB A100 or H100 PCIe holds it with room for 89,389 tokens of context in Unsloth's test, and a 48 GB card holds it only for short sequences, 12,106 tokens in the same test. The H200 NVL's 141 GB leaves the most room. 16-bit LoRA at 70B needs two GPUs, and full fine-tuning needs a cluster.

Is QLoRA worse than LoRA?

Not by much in the published evidence. The QLoRA paper reports that its 4-bit method preserved full 16-bit fine-tuning task performance in its experiments, and Unsloth says its dynamic 4-bit quants recover most of the remaining gap. Unsloth also advises training in the precision you plan to serve in, so if you will serve the model in 16-bit, LoRA on a 16-bit base is the closer match.

Not for one GPU, and rarely for LoRA or QLoRA on two. With an adapter and plain data parallel training, only the adapter's gradients cross between GPUs, and Axolotl puts the adapter at 0.1 to 1% of the parameters. Full sharding with FSDP or ZeRO stage 3 moves weights and gradients on every step, and that is where the link between GPUs starts to matter. The offers API flags the A100 SXM4 and H200 NVL offers as NVLink. Check the links on your own VM with nvidia-smi topo -m, and see InfiniBand vs NVLink for what each link does.

How much memory does gpt-oss need for fine-tuning?

Unsloth lists QLoRA at 14 GB for gpt-oss-20b and 65 GB for gpt-oss-120b, and 16-bit LoRA at 44 GB and 210 GB. The 120b QLoRA run therefore wants an 80 GB card or the H200 NVL. The gpt-oss GPU requirements guide covers serving the same models. Fine-tuning gpt-oss-20b walks through a run on one GPU.


For most readers the answer is one GPU: a 48 GB card for QLoRA up to 32B, and an 80 GB card or an H200 NVL for 70B. Run the Unsloth tutorial on an RTX A6000 first, read the peak memory it prints, then pick your GPU from the catalog with that number in hand. For inference, image and video workloads, the VRAM guide has the same kind of table.

Keep building

Choose your next step.