GPU guide

Fine-tune an LLM with Unsloth on a cloud GPU (QLoRA, step by step)

Unsloth QLoRA fine-tuning on one RTX A6000: install, train Gemma 4 E4B, push the LoRA adapter to Hugging Face before the VM stops, and size bigger runs.

Faiz Ahmed12 min read

Unsloth fine-tuning with QLoRA is where I would start on a single rented GPU: Unsloth loads the base model in 4-bit, trains a small LoRA adapter on top, and puts the memory for Gemma 4 E4B at about 10 GB. A 48 GB RTX A6000 has room to spare. The run below goes from launch to a saved adapter: start the PyTorch + Jupyter template on a QuantaCloud GPU instance, install Unsloth in its own environment, train on 3,000 chat examples for 60 steps, push the adapter to Hugging Face, and stop the instance. The table at the end sizes bigger models.

ItemThis run
GPURTX A6000, 48 GB, 1 GPU, $0.48/GPU-hr
TemplatePyTorch + Jupyter (the Bare Metal template works too)
ModelGemma 4 E4B instruct, unsloth/gemma-4-E4B-it, Apache-2.0
MethodQLoRA: 4-bit base, LoRA rank 8 on the attention and MLP layers
DataThe first 3,000 rows of mlabonne/FineTome-100k
Training60 steps, batch size 1, gradient accumulation 4, 2,048-token sequences

What you need#

You need a QuantaCloud account with credits, a Hugging Face account, and a Hugging Face access token with write permission. The minimum deposit is $5, and your balance has to cover at least one hour of the GPU to launch it (adding credits). Gemma 4 is released under Apache-2.0 and is not gated on Hugging Face, so there is no license form to accept first.

Launch the GPU#

The RTX A6000 is my pick for this run: 48 GB is far more than E4B's 10 GB, and it was $0.48 per GPU-hour on 2026-09-27 (now $0.48/GPU-hr).

  1. Open the deploy page with the offer and the PyTorch + Jupyter template selected: Launch an RTX A6000 with PyTorch + Jupyter
  2. Click Deploy. If you add credit on the way, check that the template still reads PyTorch + Jupyter before you click. The deploy docs walk through the page.
  3. When the deployment shows Running, click Open Application on the Deployments page. JupyterLab opens at a private URL behind your QuantaCloud login, and only your account can open it (templates docs).
  4. In JupyterLab, open a terminal: File, New, Terminal.

If you prefer SSH, launch with the Bare Metal template instead (RTX A6000 with Bare Metal ), connect as ubuntu (SSH docs, or our guide to SSH and VS Code), and run the same commands in that shell.

Check the GPU, the driver and a compiler#

Two checks save a failed install later: the driver version and a C compiler.

Terminal
nvidia-smi
gcc --version

nvidia-smi should list one NVIDIA RTX A6000 with about 48 GB and no running processes. Note the driver version in its header, because the next step installs a PyTorch build to match it. The compiler matters because Triton, which Unsloth runs on, builds small helpers with a C compiler the first time it runs. If gcc is missing on Bare Metal, install it with sudo apt-get update && sudo apt-get install -y build-essential, the package Unsloth lists for Linux.

Install Unsloth in its own environment#

The one thing I always do is give Unsloth its own environment. Unsloth 2026.9.11 requires torch below 2.13 and transformers up to 5.5.0, while vLLM 0.30.0 pins torch 2.13.0, so these two releases cannot share one. A separate environment also leaves the template's own PyTorch untouched.

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/unsloth-env --python 3.13
source ~/unsloth-env/bin/activate
uv pip install unsloth timm --torch-backend=auto

These are the install commands from Unsloth's docs, plus timm, which Unsloth's own Gemma 4 notebook installs for the model's vision and audio parts. The flag --torch-backend=auto makes uv read the installed NVIDIA driver and pick the PyTorch CUDA build that matches it. That matters because a plain install of PyTorch from PyPI now brings CUDA 13.0 wheels, which need driver 580 or newer. uv also downloads Python 3.13 itself, so the system Python on the instance does not matter.

Check that PyTorch sees the GPU, and keep the version list for later:

Terminal
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"
uv pip list | grep -Ei "^(unsloth|unsloth-zoo|torch|transformers|trl|peft|bitsandbytes) "

The first line should end with True NVIDIA RTX A6000. Every new terminal needs source ~/unsloth-env/bin/activate before it can use the environment. To run the same code from a notebook instead, register the environment as a kernel, reload JupyterLab and pick Python (unsloth) in the launcher:

Terminal
uv pip install ipykernel
python -m ipykernel install --user --name unsloth --display-name "Python (unsloth)"

Log in to Hugging Face#

The adapter goes to Hugging Face at the end, so log in now:

Terminal
hf auth login

It asks how you want to log in. Choose Paste an access token and paste a token with write permission, created under Settings, Access Tokens on huggingface.co, because the upload at the end writes to your account. The token is saved on the instance's disk, which is deleted when you stop, so it does not outlive the session.

Write the training script#

The script follows Unsloth's own Gemma 4 E4B text notebook, with four changes: 2,048-token sequences instead of 1,024, an outputs folder for checkpoints, a named folder for the adapter, and memory counters at the end. Create train.py in your home folder (in JupyterLab: File, New, Python File, then rename it) and paste this in:

Python
# train.py: QLoRA fine-tune of Gemma 4 E4B with Unsloth on one GPU
from unsloth import FastModel  # import Unsloth before trl and transformers
from unsloth.chat_templates import (
    get_chat_template,
    standardize_data_formats,
    train_on_responses_only,
)
import torch
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

MODEL = "unsloth/gemma-4-E4B-it"
OUT_DIR = "gemma4-e4b-lora"
MAX_SEQ_LENGTH = 2048

model, tokenizer = FastModel.from_pretrained(
    model_name = MODEL,
    dtype = None,                   # None = auto-detect
    max_seq_length = MAX_SEQ_LENGTH,
    load_in_4bit = True,            # QLoRA: 4-bit base weights
    full_finetuning = False,
)

model = FastModel.get_peft_model(
    model,
    finetune_vision_layers = False,     # text-only data
    finetune_language_layers = True,
    finetune_attention_modules = True,
    finetune_mlp_modules = True,
    r = 8,
    lora_alpha = 8,
    lora_dropout = 0,
    bias = "none",
    random_state = 3407,
)

tokenizer = get_chat_template(tokenizer, chat_template = "gemma-4")

dataset = load_dataset("mlabonne/FineTome-100k", split = "train[:3000]")
dataset = standardize_data_formats(dataset)

def to_text(batch):
    texts = [
        tokenizer.apply_chat_template(
            convo, tokenize = False, add_generation_prompt = False
        ).removeprefix("<bos>")
        for convo in batch["conversations"]
    ]
    return {"text": texts}

dataset = dataset.map(to_text, batched = True)

trainer = SFTTrainer(
    model = model,
    tokenizer = tokenizer,
    train_dataset = dataset,
    args = SFTConfig(
        dataset_text_field = "text",
        per_device_train_batch_size = 1,
        gradient_accumulation_steps = 4,
        warmup_steps = 5,
        max_steps = 60,                 # a test run; use num_train_epochs = 1 for a real one
        learning_rate = 2e-4,
        logging_steps = 1,
        optim = "adamw_8bit",
        weight_decay = 0.001,
        lr_scheduler_type = "linear",
        seed = 3407,
        output_dir = "outputs",
        report_to = "none",
    ),
)
trainer = train_on_responses_only(trainer)

torch.cuda.reset_peak_memory_stats()
stats = trainer.train()

gib = 2**30
print(f"train_runtime_s = {stats.metrics['train_runtime']:.0f}")
print(f"peak_allocated_gib = {torch.cuda.max_memory_allocated() / gib:.2f}")
print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved() / gib:.2f}")

model.save_pretrained(OUT_DIR)
tokenizer.save_pretrained(OUT_DIR)
print(f"adapter saved to {OUT_DIR}")

# Smoke test: one short answer from the fine-tuned model
messages = [{"role": "user", "content": [
    {"type": "text", "text": "Continue the sequence: 1, 1, 2, 3, 5, 8,"},
]}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt = True,
    tokenize = True,
    return_dict = True,
    return_tensors = "pt",
).to("cuda")
outputs = model.generate(**inputs, max_new_tokens = 64, temperature = 1.0, top_p = 0.95, top_k = 64)
print(tokenizer.batch_decode(outputs)[0])

The settings that decide memory and results:

SettingValueWhy
load_in_4bitTrueQLoRA: the frozen base stays in 4-bit
r, lora_alpha8, 8The values in Unsloth's notebook for this model
finetune_vision_layersFalseThe data is text only
per_device_train_batch_size, gradient_accumulation_steps1, 4An effective batch of 4 at the lowest memory
max_steps60A test run. For a real run, use num_train_epochs = 1 instead
optimadamw_8bit8-bit optimizer states
train_on_responses_onlyonThe loss counts only the model's answers, not the prompts

To train on your own data instead, give the dataset the same conversations column of user and assistant turns and change the load_dataset line.

Run it and watch the memory#

Run the script in the terminal and keep the log:

Terminal
python train.py 2>&1 | tee train.log

In a second terminal, watch the GPU while it trains:

Terminal
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv -l 5

The first minutes go to downloading the model and compiling kernels. Unsloth's benchmarks note that torch.compile can take about 5 minutes or longer to warm up, so judge the speed after that. A training loss of 13 to 15 is expected here: Unsloth says that is normal for Gemma 4 E2B and E4B. When training finishes, the script prints the training time and peak memory, saves the adapter to gemma4-e4b-lora, and generates one short answer as a smoke test.

On Bare Metal, start longer runs with nohup python train.py > train.log 2>&1 & and follow them with tail -f train.log, so a dropped SSH connection does not stop the training.

Save the adapter off the VM before you stop#

Stopping a QuantaCloud instance terminates it and deletes its disk. There are no volumes or snapshots to fall back on, so the adapter has to leave the VM before you stop. Push it to a private Hugging Face repo:

Terminal
hf upload your-username/gemma4-e4b-lora ./gemma4-e4b-lora . --private

The folder holds the LoRA weights and the tokenizer files, not the base model, so it is small. The command creates the repo if it does not exist yet and prints its URL.

To keep a copy on your own machine as well, pack the folder and download it from the JupyterLab file browser (right-click, Download):

Terminal
tar czf gemma4-e4b-lora.tgz gemma4-e4b-lora

On Bare Metal, copy it from your own computer with scp:

Terminal
scp -r ubuntu@<instance-ip>:~/gemma4-e4b-lora .

For long runs, do not wait for the end. Save checkpoints along the way (save_steps in SFTConfig) and push the output folder on a schedule from a second terminal. This uploads it every 15 minutes until you press Ctrl+C:

Terminal
hf upload your-username/gemma4-e4b-run ./outputs . --private --every=15

Load the adapter in a later session#

A later session needs only the adapter and the same environment. On a new instance, repeat the install and login steps, then load the adapter by its repo name, the same way Unsloth's notebook loads a saved folder:

Python
from unsloth import FastModel

model, tokenizer = FastModel.from_pretrained(
    model_name = "your-username/gemma4-e4b-lora",
    max_seq_length = 2048,
    load_in_4bit = True,
)

Export it for serving#

An adapter is enough to keep training or to test the model, but serving stacks want one model folder. Add one of these lines to the end of train.py, or run it after loading the adapter as above. For vLLM:

Python
model.save_pretrained_merged("gemma4-e4b-merged", tokenizer, save_method = "merged_16bit")

That merges the adapter into the base weights in 16-bit, about 16 GB for this model, the size of Google's BF16 checkpoint on the Hub (15.99 GB). Serve it from a separate vLLM environment or container, since vLLM pins a different PyTorch: Deploy vLLM with Docker and Install vLLM cover both. For llama.cpp and Ollama, write a GGUF file instead:

Python
model.save_pretrained_gguf("gemma4-e4b-gguf", tokenizer, quantization_method = "Q8_0")

Unsloth's Gemma 4 notebook lists Q8_0, BF16 and F16 as the GGUF types it supports for this model for now. Push either folder off the VM with hf upload before you stop, like the adapter.

Stop the instance and what it costs#

Stop the instance from the Deployments page once the adapter is safe. Stopping terminates it and deletes the disk. The first hour was charged at launch, each further hour is charged when the previous one is used up, and the unused seconds of the current hour are refunded when you stop, so a session costs its duration times the hourly price. For example, a 45-minute session on an RTX A6000 at $0.48/hr on 2026-09-27 cost $0.36 (our calculation: 0.75 x $0.48). Today's price is $0.48/GPU-hr.

For a run of several hours, deposit enough first or turn on auto top-up: if your balance cannot cover the next hour, the instance is terminated and its disk deleted with the checkpoints on it. The billing docs and pricing have the details.

Bigger models: what fits where#

The same script scales by changing the model name and, past 48 GB, the GPU. Unsloth's published minimums, at short sequence lengths, with the GPU I would use:

ModelQLoRA (4-bit)LoRA (16-bit)Where I would run it
Gemma 4 E4B (this guide)10 GB17 GBAny 48 GB GPU: RTX A6000, RTX 6000 Ada, L40, L40S
14B class8.5 GB33 GBAny 48 GB GPU
Gemma 4 31B22 GBnot publishedA 48 GB GPU for QLoRA
Qwen3.8-27B24 GBmore than 36 GBQLoRA on 48 GB, LoRA on an 80 GB A100 or H100 PCIe
32B class26 GB76 GBQLoRA on 48 GB, LoRA on the 96 GB RTX PRO 6000 or the 141 GB H200 NVL
gpt-oss-120b65 GB210 GBQLoRA on an 80 GB GPU or the H200 NVL
70B class41 GB164 GBQLoRA on an 80 GB GPU or the H200 NVL, LoRA on 2x H200 NVL

These are floors, not comfortable fits: longer sequences and bigger batches add activations on top. The LoRA, QLoRA and full fine-tuning VRAM guide compares Unsloth's figures with Axolotl's and LLaMA-Factory's and shows how context length moves them, and the CUDA out of memory guide covers what to try when a run does not fit. For gpt-oss, fine-tuning gpt-oss-20b covers a run on a single GPU.

Every QLoRA row fits one GPU, and so does 16-bit LoRA up to the 32B class. For two or more GPUs, Unsloth documents data parallel training with torchrun --nproc_per_node 2 train.py or accelerate launch train.py, and device_map = "balanced" to split a model that does not fit one GPU. QuantaCloud's 2x, 4x and 8x VMs run the same commands.


My rule for a first run: take the smallest GPU that holds the model with room to spare, read the peak memory train.py prints, and only then size up. For Gemma 4 E4B, and for QLoRA of anything up to 32B, that is a 48 GB card like the RTX A6000. For 70B QLoRA it is an 80 GB card or the H200 NVL.

Launch an RTX A6000 with PyTorch + Jupyter

Keep building

Choose your next step.