Unsloth fine-tuning with QLoRA is where I would start on a single rented GPU: Unsloth loads the base model in 4-bit, trains a small LoRA adapter on top, and puts the memory for Gemma 4 E4B at about 10 GB. A 48 GB RTX A6000 has room to spare. The run below goes from launch to a saved adapter: start the PyTorch + Jupyter template on a QuantaCloud GPU instance, install Unsloth in its own environment, train on 3,000 chat examples for 60 steps, push the adapter to Hugging Face, and stop the instance. The table at the end sizes bigger models.
| Item | This run |
|---|---|
| GPU | RTX A6000, 48 GB, 1 GPU, $0.48/GPU-hr |
| Template | PyTorch + Jupyter (the Bare Metal template works too) |
| Model | Gemma 4 E4B instruct, unsloth/gemma-4-E4B-it, Apache-2.0 |
| Method | QLoRA: 4-bit base, LoRA rank 8 on the attention and MLP layers |
| Data | The first 3,000 rows of mlabonne/FineTome-100k |
| Training | 60 steps, batch size 1, gradient accumulation 4, 2,048-token sequences |
What you need#
You need a QuantaCloud account with credits, a Hugging Face account, and a Hugging Face access token with write permission. The minimum deposit is $5, and your balance has to cover at least one hour of the GPU to launch it (adding credits). Gemma 4 is released under Apache-2.0 and is not gated on Hugging Face, so there is no license form to accept first.
Launch the GPU#
The RTX A6000 is my pick for this run: 48 GB is far more than E4B's 10 GB, and it was $0.48 per GPU-hour on 2026-09-27 (now $0.48/GPU-hr).
- Open the deploy page with the offer and the PyTorch + Jupyter template selected: Launch an RTX A6000 with PyTorch + Jupyter
- Click Deploy. If you add credit on the way, check that the template still reads PyTorch + Jupyter before you click. The deploy docs walk through the page.
- When the deployment shows Running, click Open Application on the Deployments page. JupyterLab opens at a private URL behind your QuantaCloud login, and only your account can open it (templates docs).
- In JupyterLab, open a terminal: File, New, Terminal.
If you prefer SSH, launch with the Bare Metal template instead
(RTX A6000 with Bare Metal ),
connect as ubuntu (SSH docs, or our guide to
SSH and VS Code), and run the same commands in that shell.
Check the GPU, the driver and a compiler#
Two checks save a failed install later: the driver version and a C compiler.
nvidia-smi
gcc --version
nvidia-smi should list one NVIDIA RTX A6000 with about 48 GB and no running processes. Note the driver version in
its header, because the next step installs a PyTorch build to match it. The compiler matters because Triton, which
Unsloth runs on, builds small helpers with a C compiler the first time it runs. If gcc is missing on Bare Metal,
install it with sudo apt-get update && sudo apt-get install -y build-essential, the package Unsloth lists for Linux.
Install Unsloth in its own environment#
The one thing I always do is give Unsloth its own environment. Unsloth 2026.9.11 requires torch below 2.13 and transformers up to 5.5.0, while vLLM 0.30.0 pins torch 2.13.0, so these two releases cannot share one. A separate environment also leaves the template's own PyTorch untouched.
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/unsloth-env --python 3.13
source ~/unsloth-env/bin/activate
uv pip install unsloth timm --torch-backend=auto
These are the install commands from Unsloth's docs, plus
timm, which Unsloth's own Gemma 4 notebook installs for the model's vision and audio parts. The flag
--torch-backend=auto makes uv read the installed NVIDIA driver and pick the PyTorch CUDA build that matches it. That
matters because a plain install of PyTorch from PyPI now brings CUDA 13.0 wheels, which need driver 580 or newer. uv
also downloads Python 3.13 itself, so the system Python on the instance does not matter.
Check that PyTorch sees the GPU, and keep the version list for later:
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"
uv pip list | grep -Ei "^(unsloth|unsloth-zoo|torch|transformers|trl|peft|bitsandbytes) "
The first line should end with True NVIDIA RTX A6000. Every new terminal needs source ~/unsloth-env/bin/activate
before it can use the environment. To run the same code from a notebook instead, register the environment as a
kernel, reload JupyterLab and pick Python (unsloth) in the launcher:
uv pip install ipykernel
python -m ipykernel install --user --name unsloth --display-name "Python (unsloth)"
Log in to Hugging Face#
The adapter goes to Hugging Face at the end, so log in now:
hf auth login
It asks how you want to log in. Choose Paste an access token and paste a token with write permission, created under Settings, Access Tokens on huggingface.co, because the upload at the end writes to your account. The token is saved on the instance's disk, which is deleted when you stop, so it does not outlive the session.
Write the training script#
The script follows Unsloth's own Gemma 4 E4B text notebook, with four changes: 2,048-token sequences instead of
1,024, an outputs folder for checkpoints, a named folder for the adapter, and memory counters at the end. Create train.py in your home folder (in
JupyterLab: File, New, Python File, then rename it) and paste this in:
# train.py: QLoRA fine-tune of Gemma 4 E4B with Unsloth on one GPU
from unsloth import FastModel # import Unsloth before trl and transformers
from unsloth.chat_templates import (
get_chat_template,
standardize_data_formats,
train_on_responses_only,
)
import torch
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer
MODEL = "unsloth/gemma-4-E4B-it"
OUT_DIR = "gemma4-e4b-lora"
MAX_SEQ_LENGTH = 2048
model, tokenizer = FastModel.from_pretrained(
model_name = MODEL,
dtype = None, # None = auto-detect
max_seq_length = MAX_SEQ_LENGTH,
load_in_4bit = True, # QLoRA: 4-bit base weights
full_finetuning = False,
)
model = FastModel.get_peft_model(
model,
finetune_vision_layers = False, # text-only data
finetune_language_layers = True,
finetune_attention_modules = True,
finetune_mlp_modules = True,
r = 8,
lora_alpha = 8,
lora_dropout = 0,
bias = "none",
random_state = 3407,
)
tokenizer = get_chat_template(tokenizer, chat_template = "gemma-4")
dataset = load_dataset("mlabonne/FineTome-100k", split = "train[:3000]")
dataset = standardize_data_formats(dataset)
def to_text(batch):
texts = [
tokenizer.apply_chat_template(
convo, tokenize = False, add_generation_prompt = False
).removeprefix("<bos>")
for convo in batch["conversations"]
]
return {"text": texts}
dataset = dataset.map(to_text, batched = True)
trainer = SFTTrainer(
model = model,
tokenizer = tokenizer,
train_dataset = dataset,
args = SFTConfig(
dataset_text_field = "text",
per_device_train_batch_size = 1,
gradient_accumulation_steps = 4,
warmup_steps = 5,
max_steps = 60, # a test run; use num_train_epochs = 1 for a real one
learning_rate = 2e-4,
logging_steps = 1,
optim = "adamw_8bit",
weight_decay = 0.001,
lr_scheduler_type = "linear",
seed = 3407,
output_dir = "outputs",
report_to = "none",
),
)
trainer = train_on_responses_only(trainer)
torch.cuda.reset_peak_memory_stats()
stats = trainer.train()
gib = 2**30
print(f"train_runtime_s = {stats.metrics['train_runtime']:.0f}")
print(f"peak_allocated_gib = {torch.cuda.max_memory_allocated() / gib:.2f}")
print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved() / gib:.2f}")
model.save_pretrained(OUT_DIR)
tokenizer.save_pretrained(OUT_DIR)
print(f"adapter saved to {OUT_DIR}")
# Smoke test: one short answer from the fine-tuned model
messages = [{"role": "user", "content": [
{"type": "text", "text": "Continue the sequence: 1, 1, 2, 3, 5, 8,"},
]}]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt = True,
tokenize = True,
return_dict = True,
return_tensors = "pt",
).to("cuda")
outputs = model.generate(**inputs, max_new_tokens = 64, temperature = 1.0, top_p = 0.95, top_k = 64)
print(tokenizer.batch_decode(outputs)[0])
The settings that decide memory and results:
| Setting | Value | Why |
|---|---|---|
load_in_4bit | True | QLoRA: the frozen base stays in 4-bit |
r, lora_alpha | 8, 8 | The values in Unsloth's notebook for this model |
finetune_vision_layers | False | The data is text only |
per_device_train_batch_size, gradient_accumulation_steps | 1, 4 | An effective batch of 4 at the lowest memory |
max_steps | 60 | A test run. For a real run, use num_train_epochs = 1 instead |
optim | adamw_8bit | 8-bit optimizer states |
train_on_responses_only | on | The loss counts only the model's answers, not the prompts |
To train on your own data instead, give the dataset the same conversations column of user and assistant turns and
change the load_dataset line.
Run it and watch the memory#
Run the script in the terminal and keep the log:
python train.py 2>&1 | tee train.log
In a second terminal, watch the GPU while it trains:
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv -l 5
The first minutes go to downloading the model and compiling kernels. Unsloth's benchmarks note that torch.compile
can take about 5 minutes or longer to warm up, so judge the speed after that. A training loss of 13 to 15 is expected
here: Unsloth says that is normal for Gemma 4 E2B and E4B. When training finishes, the script prints the training
time and peak memory, saves the adapter to gemma4-e4b-lora, and generates one short answer as a smoke test.
On Bare Metal, start longer runs with nohup python train.py > train.log 2>&1 & and follow them with
tail -f train.log, so a dropped SSH connection does not stop the training.
Save the adapter off the VM before you stop#
Stopping a QuantaCloud instance terminates it and deletes its disk. There are no volumes or snapshots to fall back on, so the adapter has to leave the VM before you stop. Push it to a private Hugging Face repo:
hf upload your-username/gemma4-e4b-lora ./gemma4-e4b-lora . --private
The folder holds the LoRA weights and the tokenizer files, not the base model, so it is small. The command creates the repo if it does not exist yet and prints its URL.
To keep a copy on your own machine as well, pack the folder and download it from the JupyterLab file browser (right-click, Download):
tar czf gemma4-e4b-lora.tgz gemma4-e4b-lora
On Bare Metal, copy it from your own computer with scp:
scp -r ubuntu@<instance-ip>:~/gemma4-e4b-lora .
For long runs, do not wait for the end. Save checkpoints along the way (save_steps in SFTConfig) and push the
output folder on a schedule from a second terminal. This uploads it every 15 minutes until you press Ctrl+C:
hf upload your-username/gemma4-e4b-run ./outputs . --private --every=15
Load the adapter in a later session#
A later session needs only the adapter and the same environment. On a new instance, repeat the install and login steps, then load the adapter by its repo name, the same way Unsloth's notebook loads a saved folder:
from unsloth import FastModel
model, tokenizer = FastModel.from_pretrained(
model_name = "your-username/gemma4-e4b-lora",
max_seq_length = 2048,
load_in_4bit = True,
)
Export it for serving#
An adapter is enough to keep training or to test the model, but serving stacks want one model folder. Add one of
these lines to the end of train.py, or run it after loading the adapter as above. For vLLM:
model.save_pretrained_merged("gemma4-e4b-merged", tokenizer, save_method = "merged_16bit")
That merges the adapter into the base weights in 16-bit, about 16 GB for this model, the size of Google's BF16 checkpoint on the Hub (15.99 GB). Serve it from a separate vLLM environment or container, since vLLM pins a different PyTorch: Deploy vLLM with Docker and Install vLLM cover both. For llama.cpp and Ollama, write a GGUF file instead:
model.save_pretrained_gguf("gemma4-e4b-gguf", tokenizer, quantization_method = "Q8_0")
Unsloth's Gemma 4 notebook lists Q8_0, BF16 and F16 as the GGUF types it supports for this model for now. Push either
folder off the VM with hf upload before you stop, like the adapter.
Stop the instance and what it costs#
Stop the instance from the Deployments page once the adapter is safe. Stopping terminates it and deletes the disk. The first hour was charged at launch, each further hour is charged when the previous one is used up, and the unused seconds of the current hour are refunded when you stop, so a session costs its duration times the hourly price. For example, a 45-minute session on an RTX A6000 at $0.48/hr on 2026-09-27 cost $0.36 (our calculation: 0.75 x $0.48). Today's price is $0.48/GPU-hr.
For a run of several hours, deposit enough first or turn on auto top-up: if your balance cannot cover the next hour, the instance is terminated and its disk deleted with the checkpoints on it. The billing docs and pricing have the details.
Bigger models: what fits where#
The same script scales by changing the model name and, past 48 GB, the GPU. Unsloth's published minimums, at short sequence lengths, with the GPU I would use:
| Model | QLoRA (4-bit) | LoRA (16-bit) | Where I would run it |
|---|---|---|---|
| Gemma 4 E4B (this guide) | 10 GB | 17 GB | Any 48 GB GPU: RTX A6000, RTX 6000 Ada, L40, L40S |
| 14B class | 8.5 GB | 33 GB | Any 48 GB GPU |
| Gemma 4 31B | 22 GB | not published | A 48 GB GPU for QLoRA |
| Qwen3.8-27B | 24 GB | more than 36 GB | QLoRA on 48 GB, LoRA on an 80 GB A100 or H100 PCIe |
| 32B class | 26 GB | 76 GB | QLoRA on 48 GB, LoRA on the 96 GB RTX PRO 6000 or the 141 GB H200 NVL |
| gpt-oss-120b | 65 GB | 210 GB | QLoRA on an 80 GB GPU or the H200 NVL |
| 70B class | 41 GB | 164 GB | QLoRA on an 80 GB GPU or the H200 NVL, LoRA on 2x H200 NVL |
These are floors, not comfortable fits: longer sequences and bigger batches add activations on top. The LoRA, QLoRA and full fine-tuning VRAM guide compares Unsloth's figures with Axolotl's and LLaMA-Factory's and shows how context length moves them, and the CUDA out of memory guide covers what to try when a run does not fit. For gpt-oss, fine-tuning gpt-oss-20b covers a run on a single GPU.
Every QLoRA row fits one GPU, and so does 16-bit LoRA up to the 32B class. For two or more GPUs, Unsloth documents
data parallel training with torchrun --nproc_per_node 2 train.py or accelerate launch train.py, and
device_map = "balanced" to split a model that does not fit one GPU. QuantaCloud's 2x, 4x and 8x VMs run the same
commands.
My rule for a first run: take the smallest GPU that holds the model with room to spare, read the peak memory
train.py prints, and only then size up. For Gemma 4 E4B, and for QLoRA of anything up to 32B, that is a 48 GB card
like the RTX A6000. For 70B QLoRA it is an 80 GB card or the H200 NVL.