GPU guide

Multi-GPU fine-tuning with Axolotl on one 8-GPU VM

Install Axolotl, write a LoRA config, spread it over 8 GPUs with FSDP2 or DeepSpeed, size jobs with Axolotl's table, and push the adapter off the VM.

Faiz Ahmed11 min read

Axolotl turns a fine-tune into one YAML file and one command, axolotl train config.yml, and on a multi-GPU VM that command already starts one process per GPU. Adding fsdp_version: 2 and an fsdp_config block, or a single deepspeed: line, makes the same file shard the model across all eight GPUs. The config below trains a LoRA adapter on Qwen3-32B, which Axolotl's own table puts at 64 to 80 GB for LoRA. That is too much for one 48 GB card, but FSDP spreads the frozen weights to 8.2 GB per GPU across eight (our calculation: 65.5 GB / 8), and the adapter goes to the Hugging Face Hub at every save.

What fits: Axolotl's memory table#

Axolotl's own table is the one I plan with, because it states its conditions: a context of 512 to 2,048 tokens and a micro-batch of 1 to 2. Longer sequences and bigger batches need more, and gradient checkpointing gets some of it back at about 30% slower training.

Model sizeQLoRA (4-bit)LoRA (bf16)Full (bf16 + AdamW)
1-3B6-8 GB8-12 GB24-32 GB
7-8B10-14 GB16-24 GB60-80 GB
13-14B16-20 GB28-40 GB120+ GB
30-34B24-32 GB64-80 GB2-4x 80 GB
70-72B40-48 GB2x 80 GB4-8x 80 GB

Mapped onto QuantaCloud, where single GPUs have 48, 80, 96 or 141 GB and VMs go up to eight GPUs (our mapping, from the catalog on 2026-09-27):

JobWhere it runs
QLoRA up to 32B, LoRA up to 14BOne 48 GB GPU. Extra GPUs only add speed, through DDP
LoRA at 30-34BOne 96 GB RTX PRO 6000 or 141 GB H200 NVL, or sharded with FSDP over 2 to 8 GPUs of 48 GB
QLoRA at 70BOne 80 GB A100 or H100 PCIe, or the H200 NVL for long sequences
LoRA at 70BSharded over 2x H200 NVL, 4x A100 80GB or eight 48 GB GPUs
Full fine-tune at 7-8BSharded over two or more GPUs, such as 2x RTX PRO 6000 or 2x H200 NVL
Full fine-tune at 70BA cluster. Axolotl's 4-8x 80 GB assumes pure bf16 training, and with fp32 master weights the model states alone are 1.1 to 1.3 TB

The VRAM guide compares Axolotl's table with Unsloth's and LLaMA-Factory's and explains the gaps.

Pick the VM#

An 8-GPU VM is the natural home for Axolotl's multi-GPU modes. On 2026-09-27 there were three, all with Ubuntu 22.04, the NVIDIA driver and Docker on the Bare Metal template:

ConfigurationGPU memoryRAMDisk
8x A100 SXM4 80GB640 GB800 GB5,000 GB
8x L40S384 GB576 GB5,000 GB
8x RTX 6000 Ada384 GB640 GB2,800 GB

The 8x A100 SXM4 is the one the API flags as NVLink, and nvidia-smi topo -m on the VM shows whether the GPUs really share it. The L40S and RTX 6000 Ada have no NVLink, so FSDP traffic crosses PCIe. The DDP and FSDP guide works out what that traffic costs per step and gives an all-reduce test to run on the VM. Live prices per GPU-hour:

GPUMemoryFromAvailable now
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes

Prices checked 5 Oct 2026, 21:37 UTC

Launch 8x A100 SXM4 Launch 8x L40S

QuantaCloud has no Axolotl template, so launch Bare Metal and install it, or run Axolotl's Docker image. Eight-GPU VMs take about 10 minutes to reach running (median). Connect as ubuntu (SSH docs).

Install Axolotl#

The one thing I always check first is the driver, because it decides which PyTorch build Axolotl can use. Run nvidia-smi: CUDA 13 builds need driver 580 or newer, and older drivers take the CUDA 12.8 build. Then follow Axolotl's own uv install in a fresh environment:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
export UV_TORCH_BACKEND=cu130   # use cu128 if the driver is older than 580
uv venv ~/axolotl-env --python 3.12
source ~/axolotl-env/bin/activate
uv pip install setuptools
uv pip install --no-build-isolation "axolotl[deepspeed]==0.19.0"
axolotl fetch deepspeed_configs

The setuptools line is the one addition to Axolotl's own instructions. With --no-build-isolation, packages that build from source use the environment's tools, a fresh uv environment has no setuptools, and without it the install stops at rouge-score, which Axolotl pulls in through lm-eval. Keep Axolotl in its own environment. Version 0.19.0, released on 2026-09-10, pins Transformers 5.16.1, TRL 1.9.0, PEFT 0.20.0 and Accelerate 1.13.0, and its deepspeed extra installs DeepSpeed 0.18, not the current 0.19.7. It needs Python 3.11 or newer, PyTorch 2.11 or newer, and an Ampere or newer GPU for bf16, which covers every GPU QuantaCloud offers.

The Docker route skips the install. Axolotl's release images are CUDA 13.0 builds, so they need driver 580 or newer, and Docker needs the NVIDIA Container Toolkit to pass the GPUs through (Docker with NVIDIA GPUs covers both checks):

Terminal
docker run --gpus all --rm -it --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v ~/work:/workspace/work \
  axolotlai/axolotl:0.19.0-py3.12-cu130-2.12.1

The config#

This config trains a LoRA adapter on every linear layer of Qwen3-32B, on the first 20% of FineTome-100k, with the dataset settings from Axolotl's own Qwen3 example. Save it as qwen3-32b-lora.yml:

YAML
base_model: Qwen/Qwen3-32B
chat_template: qwen3

datasets:
  - path: mlabonne/FineTome-100k
    type: chat_template
    split: train[:20%]
    field_messages: conversations
    message_property_mappings:
      role: from
      content: value
dataset_prepared_path: last_run_prepared
val_set_size: 0.0
output_dir: ./outputs/qwen3-32b-lora

sequence_len: 4096
sample_packing: true
attn_implementation: sdpa

adapter: lora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.0
lora_target_linear: true

micro_batch_size: 1
gradient_accumulation_steps: 2
num_epochs: 1
optimizer: adamw_torch_fused
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.1
weight_decay: 0.0
bf16: auto
tf32: true

logging_steps: 1
saves_per_epoch: 2
save_total_limit: 2
hub_model_id: your-username/qwen3-32b-lora
hub_strategy: checkpoint

fsdp_version: 2
fsdp_config:
  offload_params: false
  cpu_ram_efficient_loading: true
  auto_wrap_policy: TRANSFORMER_BASED_WRAP
  transformer_layer_cls_to_wrap: Qwen3DecoderLayer
  state_dict_type: FULL_STATE_DICT
  reshard_after_forward: true
  activation_checkpointing: true

The keys that decide memory, speed and where the adapter ends up:

KeyValueWhy
adapter, lora_target_linearlora, true16-bit LoRA on every linear layer. load_in_4bit: true with adapter: qlora turns it into QLoRA
sequence_len, sample_packing4096, truePacks short conversations into full 4,096-token rows
attn_implementationsdpaPyTorch's own attention. Axolotl keeps packed samples apart with it, and it needs no flash-attn build
micro_batch_size, gradient_accumulation_steps1, 2With eight GPUs, 16 packed rows of up to 4,096 tokens per optimizer step
fsdp_version: 2 and fsdp_configFSDP2Shards the frozen base and the adapter across all GPUs, one Qwen3DecoderLayer at a time
activation_checkpointingtrueRecomputes activations in backward instead of storing them
hub_model_id, hub_strategyyour repo, checkpointPushes the adapter at every save, plus the newest checkpoint in a last-checkpoint folder you can resume from

The 20% slice is 20,000 conversations and about 11.1 million tokens with the Qwen3 chat template, by our count. By default Axolotl drops the 61 conversations longer than sequence_len instead of cutting them, which leaves 10.8 million tokens. At 65,536 tokens per optimizer step, one epoch is about 165 steps if packing were perfect (our calculation: 10,829,639 / 65,536), a little more in practice.

One warning appears at start-up: Axolotl 0.19.0 says sample packing with sdpa in bf16 may train to a loss of 0.0 on GPUs older than Hopper, which includes the A100 and L40S. Check that the loss in the first steps is well above zero. If it reads 0.0, set sample_packing: false, which trains one conversation per row in more, shorter steps.

FSDP2 or DeepSpeed#

Axolotl runs one strategy at a time: DDP when the config names neither, FSDP when it has fsdp_config, and DeepSpeed when it has a deepspeed: line. Its docs recommend FSDP2 for new users and mark FSDP1 as deprecated, which is why the config above uses it. To switch to DeepSpeed, delete the fsdp_version and fsdp_config lines and add one line:

YAML
deepspeed: deepspeed_configs/zero3_bf16.json

Axolotl's advice is to try stage 1, then 2, then 3, and stop at the first that fits. One detail in its bundled files matters on a fresh VM: zero2.json offloads the optimizer to CPU, and DeepSpeed's CPU optimizer compiles against the CUDA toolkit the first time it runs, while zero1.json and zero3_bf16.json keep everything on the GPUs. The DeepSpeed ZeRO guide has the memory per stage. For a LoRA job small enough for one GPU, drop both blocks and let DDP run a copy on every GPU, which is the simplest way to put all eight to work.

Run it#

Log in to Hugging Face so hub_model_id can push, tokenize the dataset once, then train:

Terminal
hf auth login
CUDA_VISIBLE_DEVICES="0" axolotl preprocess qwen3-32b-lora.yml
axolotl train qwen3-32b-lora.yml

axolotl train launches through Accelerate by default, and Accelerate starts one process per visible GPU when it has no config of its own. To use torchrun instead, pass its arguments after --: axolotl train qwen3-32b-lora.yml --launcher torchrun -- --nproc_per_node=8. For a timing run before the full epoch, add --max-steps 50. Axolotl logs the loss and memory at every step. Start long runs under nohup or tmux so a dropped SSH session does not stop them, and watch the GPUs with nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv -l 5. If a collective times out after 30 minutes, Axolotl's NCCL guide suggests NCCL_DEBUG=INFO and an NCCL bandwidth test, and ddp_timeout raises the limit.

Save the adapter off the VM#

Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots. With hub_model_id set, Axolotl turns on the Trainer's Hub upload to a private repo: each save pushes the adapter, and hub_strategy: checkpoint also pushes the newest checkpoint to a last-checkpoint folder. The adapter is small, because it holds only the LoRA weights. Everything is also written to output_dir:

Terminal
ls outputs/qwen3-32b-lora
# adapter_config.json, adapter_model.safetensors, tokenizer files, checkpoint folders

Without hub_model_id, push the folder yourself before you stop: hf upload your-username/qwen3-32b-lora ./outputs/qwen3-32b-lora . --private. To resume on a new VM, download last-checkpoint from the repo and pass it to axolotl train qwen3-32b-lora.yml --resume-from-checkpoint <folder>. For serving engines that want one set of weights, merge the adapter into the base model:

Terminal
axolotl merge-lora qwen3-32b-lora.yml --lora-model-dir="./outputs/qwen3-32b-lora"

The merged model is written to outputs/qwen3-32b-lora/merged in 16-bit, about the size of the base checkpoint, 65.5 GB for Qwen3-32B. To merge on the CPU and leave the GPUs free, prefix the command with CUDA_VISIBLE_DEVICES="": the 8-GPU VMs have 576 to 800 GB of RAM. Push the merged folder off the VM the same way, or copy it to your own storage with the file transfer guide.

Before a long run, check the balance: the first hour is charged at launch and each further hour when the previous one is used up, and if the balance cannot cover the next hour, the instance is terminated with its disk. Pricing has the rules, and the cost guide turns tokens and speed into dollars.

FAQ#

Does Axolotl use all the GPUs automatically?

Yes. axolotl train launches through Accelerate, which starts one process per visible GPU when no Accelerate config says otherwise. Without an fsdp_config or deepspeed line that means DDP, a full copy of the model on each GPU. Use CUDA_VISIBLE_DEVICES to limit it.

Should I use FSDP or DeepSpeed with Axolotl?

FSDP2, unless you need DeepSpeed's CPU or NVMe offload. Axolotl's docs recommend FSDP2 for new users, and it is part of PyTorch, so there is nothing extra to compile.

Can I fine-tune gpt-oss with Axolotl?

Yes, Axolotl ships gpt-oss examples, and its gpt-oss README lists LoRA of gpt-oss-20b at about 44 GiB on one 48 GB GPU. The gpt-oss-20b fine-tuning guide walks through the TRL recipe instead.

What is the difference between Axolotl and Unsloth?

Axolotl is a config-driven wrapper around the Hugging Face stack with DDP, DeepSpeed and FSDP built in, which makes it strong on multi-GPU VMs. Unsloth is built around one GPU, with its own kernels to cut memory and time, and the Unsloth guide covers that path.


My rule for Axolotl on one VM: if the job fits one GPU in Axolotl's table, run it on one and use extra GPUs only for DDP speed. If it does not, add the FSDP2 block and rent an 8-GPU VM: the 8x A100 SXM4, flagged NVLink, when the traffic between GPUs matters most and nvidia-smi topo -m confirms the links, the 8x L40S or 8x RTX 6000 Ada when the price does. Full fine-tunes past one VM are cluster jobs, which we order and build to your spec as GPU clusters: send a capacity brief. The fine-tuning overview compares Axolotl with the other ways to train.

Launch 8x A100 SXM4

Keep building

Choose your next step.