GPU guide

Multi-GPU training on one node with PyTorch DDP and FSDP

DDP or FSDP2 on one 2 to 8 GPU VM: memory per GPU, a runnable torchrun script, communication without NVLink, and checkpoints that survive the VM.

Faiz Ahmed17 min read

Use DDP when one GPU can hold the whole model with its gradients and optimizer state, and FSDP when it cannot. Both launch the same way on one multi-GPU VM, one process per GPU with torchrun --standalone --nproc-per-node=gpu, and the script below runs either with one flag. The difference is memory: a full fine-tune of Qwen3-1.7B holds 32.5 GB of weights, gradients and AdamW state on every GPU under DDP, and 4.1 GB per GPU under FSDP across eight (our calculation: 2.03 billion parameters x 16 bytes, then divided by 8). The price FSDP pays is communication: it gathers each layer's weights from the other GPUs before using them, on top of reducing the gradients, and that traffic is where NVLink, or its absence, starts to matter.

DDP or FSDP: the rule#

The rule I follow is to start with DDP and move to FSDP only when DDP runs out of memory, because DDP is simpler and moves no more data per step. The two differ in what each GPU keeps and what crosses between GPUs:

DDP (DistributedDataParallel)FSDP2 (fully_shard)
Each GPU holdsA full copy of the weights, gradients and optimizer state1/N of each, plus the layer it is computing
Communication per stepOne all-reduce of the gradients during backwardAn all-gather of each layer's weights in forward and again in backward, and a reduce-scatter of its gradients
Model states per GPU, AdamW with fp32 master weights16 bytes per parameter, 20 with DDP's default gradient buckets16 bytes per parameter divided by N
Largest full fine-tune, model states onlyAbout 3B parameters on a 48 GB GPU and 5B on 80 GB (our calculation: memory / 16 bytes)Up to the VM's total GPU memory / 16 bytes
Code changeWrap the model in DDP(...)Call fully_shard on each layer, then on the whole model

The 16 bytes are the ZeRO paper's accounting for mixed-precision Adam, and FSDP is the same idea as its stage 3. PyTorch's own tutorial describes FSDP as DDP's all-reduce split into a reduce-scatter and an all-gather, and the ZeRO paper puts full sharding at 1.5 times the communication volume of plain data parallelism. FSDP1, the old FullyShardedDataParallel class, is deprecated. Use FSDP2's fully_shard, as below.

Memory per GPU for a full fine-tune#

The table is our calculation of model states only: weights, gradients and AdamW state at 16 bytes per parameter. Activations come on top and grow with sequence length and batch size, which is why the script recomputes them in the backward pass.

ModelDDP, every GPUFSDP on 2 GPUsFSDP on 4 GPUsFSDP on 8 GPUs
Qwen3-1.7B (2.03B parameters)32.5 GB16.3 GB8.1 GB4.1 GB
Qwen3-8B (8.19B parameters)131.1 GB65.5 GB32.8 GB16.4 GB

Two readings matter. Qwen3-1.7B fits DDP on a 48 GB GPU with 15.5 GB left for activations (our calculation: 48 - 32.5), as long as you pass gradient_as_bucket_view=True: PyTorch's docs say it saves memory equal to the total gradient size, which is another 8.1 GB here. Qwen3-8B does not fit DDP on anything short of the 141 GB H200 NVL, and even there it leaves about 18 GiB for activations, so it is an FSDP job: 16.4 GB per GPU on eight 48 GB L40S or RTX 6000 Ada cards (our calculation: 140.4 GiB reported minus 131.1 GB, 122.1 GiB, and 131.1 / 8). For LoRA and QLoRA, where the base model is frozen, the VRAM guide has the numbers, and most of those jobs need one GPU and plain DDP at most.

Pick the multi-GPU VM#

Any QuantaCloud GPU instance with 2, 4 or 8 GPUs runs this. On 2026-09-27 the largest configurations were:

ConfigurationGPU memoryRAMLink between GPUs
8x A100 SXM4 80GB640 GB800 GBSXM, flagged NVLink by the API
8x L40S384 GB576 GBPCIe Gen4 x16, no NVLink on this card
8x RTX 6000 Ada384 GB640 GBPCIe 4.0 x16, no NVLink on this card
4x RTX PRO 6000 Blackwell384 GB576 GBPCIe, no NVLink on this card
2x H200 NVL282 GB360 GBFlagged NVLink by the API

RAM matters because every process loads the full model into CPU memory before sharding it: eight fp32 copies of Qwen3-8B are 262 GB (our calculation: 8 x 8.19 billion x 4 bytes). Live prices per GPU-hour:

GPUMemoryFromAvailable now
A100 SXM4 80GB80 GB$1.50/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 13:48 UTC

Launch 8x A100 SXM4 Launch 8x L40S

Eight-GPU VMs take about 10 minutes to reach running (median). Connect as ubuntu over SSH (SSH docs), or with VS Code Remote. The A100 page and the L40S page cover the cards themselves.

Check how the GPUs connect#

The first thing I check on a multi-GPU VM is the link matrix, because it decides how fast DDP and FSDP can exchange data:

Terminal
nvidia-smi
nvidia-smi topo -m
nvidia-smi topo -p2p p

In topo -m, an entry of NV plus a number between two GPUs means that many bonded NVLinks. PIX, PXB, PHB, NODE and SYS mean PCIe, with the path getting longer in that order, and SYS crosses between NUMA nodes. topo -p2p p prints OK for each pair of GPUs that supports peer-to-peer transfers over PCIe. When that peer-to-peer path is missing, NCCL falls back to its shared-memory transport, which copies through host memory. NCCL's troubleshooting guide notes two things that apply to any VM: virtual machines need PCI Access Control Services turned on, so the usual bare-metal fix of disabling it is not available, and a virtual PCI topology can cost performance. QuantaCloud's API flags the A100 SXM4 and H200 NVL offers as NVLink, and the matrix is how you confirm it on your own VM.

Set up the environment#

A virtual environment with PyTorch, Transformers and datasets is all the script needs. uv picks the PyTorch CUDA build that matches the VM's driver:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/dist-env --python 3.12
source ~/dist-env/bin/activate
uv pip install torch --torch-backend=auto
uv pip install transformers datasets
python -c "import torch; print(torch.__version__, torch.cuda.device_count(), torch.cuda.nccl.version())"

The last line should print the PyTorch version, the number of GPUs in the VM and the NCCL version bundled with PyTorch. The current releases are PyTorch 2.14.0 and Transformers 5.17.0. The driver and CUDA guide explains why the driver decides which PyTorch build you get.

The training script#

The script below full fine-tunes a Hugging Face model on TRL's Capybara conversations, with --mode ddp or --mode fsdp. It keeps fp32 master weights and computes in bf16, recomputes activations in the backward pass, saves a sharded checkpoint every 50 steps, resumes from the newest one, and writes a normal Hugging Face model folder at the end. Save it as train.py:

Python
# train.py: full fine-tuning of a Hugging Face model with DDP or FSDP2 on one VM
# Launch: torchrun --standalone --nproc-per-node=gpu train.py --mode fsdp
import argparse
import contextlib
import os
import time

import torch
import torch.distributed as dist
import torch.distributed.checkpoint as dcp
from datasets import load_dataset
from torch.distributed.checkpoint.state_dict import (
    StateDictOptions, get_model_state_dict, get_state_dict, set_state_dict)
from torch.distributed.checkpoint.stateful import Stateful
from torch.distributed.fsdp import MixedPrecisionPolicy, fully_shard
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.utils.data import DataLoader, DistributedSampler
from transformers import AutoModelForCausalLM, AutoTokenizer

parser = argparse.ArgumentParser()
parser.add_argument("--mode", choices=["ddp", "fsdp"], default="fsdp")
parser.add_argument("--model", default="Qwen/Qwen3-1.7B")
parser.add_argument("--seq-len", type=int, default=2048)
parser.add_argument("--steps", type=int, default=100)
parser.add_argument("--lr", type=float, default=1e-5)
parser.add_argument("--save-every", type=int, default=50)
parser.add_argument("--ckpt-dir", default="checkpoints")
args = parser.parse_args()

# torchrun starts one process per GPU and sets LOCAL_RANK, RANK and WORLD_SIZE
local_rank = int(os.environ["LOCAL_RANK"])
device = torch.device("cuda", local_rank)
torch.cuda.set_device(device)
dist.init_process_group(backend="nccl", device_id=device)
rank, world = dist.get_rank(), dist.get_world_size()

# Data: TRL's Capybara conversations, tokenized and packed into seq-len blocks
tok = AutoTokenizer.from_pretrained(args.model)
texts = [tok.apply_chat_template(m, tokenize=False)
         for m in load_dataset("trl-lib/Capybara", split="train")["messages"]]
flat = [t for ids in tok(texts, add_special_tokens=False)["input_ids"] for t in ids]
n = len(flat) // args.seq_len
blocks = torch.tensor(flat[: n * args.seq_len]).view(n, args.seq_len)
sampler = DistributedSampler(blocks, shuffle=True, seed=0)  # a different slice per rank
loader = DataLoader(blocks, batch_size=1, sampler=sampler, drop_last=True)

# Model: fp32 master weights, bf16 compute, activations recomputed in backward
model = AutoModelForCausalLM.from_pretrained(args.model, dtype=torch.float32, use_cache=False)
model.gradient_checkpointing_enable()
if args.mode == "ddp":
    # every GPU holds the whole model, its gradients and the optimizer state
    model = DDP(model.to(device), device_ids=[local_rank], gradient_as_bucket_view=True)
    amp = lambda: torch.autocast("cuda", dtype=torch.bfloat16)
else:
    # each GPU holds 1/N of them; layers are all-gathered just in time
    mp = MixedPrecisionPolicy(param_dtype=torch.bfloat16, reduce_dtype=torch.float32)
    for layer in model.model.layers:
        fully_shard(layer, mp_policy=mp)
    fully_shard(model, mp_policy=mp)
    amp = contextlib.nullcontext
optimizer = torch.optim.AdamW(model.parameters(), lr=args.lr, weight_decay=0.0)


class AppState(Stateful):  # lets DCP save and load model and optimizer shards
    def __init__(self, model, optimizer):
        self.model, self.optimizer = model, optimizer

    def state_dict(self):
        model_sd, optim_sd = get_state_dict(self.model, self.optimizer)
        return {"model": model_sd, "optim": optim_sd}

    def load_state_dict(self, sd):
        set_state_dict(self.model, self.optimizer,
                       model_state_dict=sd["model"], optim_state_dict=sd["optim"])


# Resume from the newest checkpoint folder if one exists
step = 0
if os.path.isdir(args.ckpt_dir):
    done = sorted(int(d.split("_")[1]) for d in os.listdir(args.ckpt_dir) if d.startswith("step_"))
    if done:
        step = done[-1]
        dcp.load({"app": AppState(model, optimizer)}, checkpoint_id=f"{args.ckpt_dir}/step_{step}")


def batches():
    epoch = 0
    while True:
        sampler.set_epoch(epoch)
        yield from loader
        epoch += 1


data = batches()
for _ in range(step):  # skip the batches the checkpoint already trained on
    next(data)

model.train()
torch.cuda.reset_peak_memory_stats(device)
timed_s, timed_tokens = 0.0, 0
while step < args.steps:
    t = time.perf_counter()
    x = next(data).to(device)
    with amp():
        loss = model(input_ids=x, labels=x).loss
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
    optimizer.step()
    optimizer.zero_grad()
    torch.cuda.synchronize(device)
    step += 1
    if step > 10:  # the first steps include warm-up, so they are not timed
        timed_s += time.perf_counter() - t
        timed_tokens += x.numel() * world
    if rank == 0 and step % 10 == 0:
        print(f"step {step} loss {loss.item():.3f}", flush=True)
    if step % args.save_every == 0 or step == args.steps:
        dcp.save({"app": AppState(model, optimizer)}, checkpoint_id=f"{args.ckpt_dir}/step_{step}")

if rank == 0 and timed_s > 0:
    print(f"tokens_per_s = {timed_tokens / timed_s:.0f}")
    print(f"peak_allocated_gib = {torch.cuda.max_memory_allocated(device) / 2**30:.1f}")
    print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved(device) / 2**30:.1f}")

# Final model in Hugging Face format, bf16, written by rank 0
full_sd = get_model_state_dict(model, options=StateDictOptions(full_state_dict=True, cpu_offload=True))
if rank == 0:
    final = AutoModelForCausalLM.from_pretrained(args.model, dtype=torch.bfloat16)
    final.load_state_dict(full_sd)
    final.save_pretrained(f"{args.ckpt_dir}/final")
    tok.save_pretrained(f"{args.ckpt_dir}/final")
dist.destroy_process_group()

The lines that decide memory and speed:

LineWhat it does
DistributedSampler(...)Gives each process its own share of the data. Without it, every GPU trains on the same batches
DDP(..., gradient_as_bucket_view=True)Makes the gradients views into DDP's communication buckets instead of a second copy
fully_shard(layer, ...) then fully_shard(model, ...)Shards layer by layer, bottom-up, as PyTorch's docs require, so each layer's full weights are gathered just before it runs and freed after
MixedPrecisionPolicy(param_dtype=bf16, reduce_dtype=fp32)Computes in bf16 and reduces gradients in fp32, while the optimizer updates fp32 shards
AdamW(model.parameters()) after shardingThe optimizer must see the sharded parameters, so it is built after fully_shard
gradient_checkpointing_enable()Recomputes activations in backward. Transformers 5 uses the non-reentrant mode by default
dcp.save(...)Every rank writes its own shard in parallel, so saving does not funnel through one GPU

Each step trains on one 2,048-token block per GPU, 16,384 tokens per step on eight GPUs (our calculation: 8 x 2,048). The script trains on whole packed conversations, which keeps it short. For chat fine-tuning with the prompts masked out, Axolotl trains on the assistant turns by default, and TRL, which the DeepSpeed ZeRO guide uses, does it with assistant_only_loss=True.

Launch it with torchrun#

The same script covers every case. Start with one GPU to get a baseline, then use all of them:

Terminal
# baseline: one GPU
CUDA_VISIBLE_DEVICES=0 torchrun --standalone --nproc-per-node=1 train.py --mode ddp --ckpt-dir ck-1gpu
# DDP on every GPU in the VM
torchrun --standalone --nproc-per-node=gpu train.py --mode ddp --ckpt-dir ck-ddp
# FSDP on every GPU, and a model DDP cannot hold
torchrun --standalone --nproc-per-node=gpu train.py --mode fsdp --model Qwen/Qwen3-8B --ckpt-dir ck-fsdp-8b

--standalone runs the rendezvous on the VM itself, and --nproc-per-node=gpu starts one process per visible GPU. Each run prints its loss every 10 steps and, at the end, tokens_per_s and the peak memory on rank 0. Scaling efficiency is the 8-GPU tokens per second divided by eight times the one-GPU figure. If it comes out low, low GPU utilization covers why GPUs wait on each other, and the cost guide turns tokens per second into hours and dollars. For long runs, start the job with nohup ... > train.log 2>&1 & so a dropped SSH session does not kill it, and watch the GPUs from a second terminal with nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv -l 5.

The core problem on a PCIe-only VM is that every step moves the same bytes whether the link is fast or not. With this script's settings, DDP all-reduces fp32 gradients, and each GPU sends and receives 2 x (N-1)/N times their size. FSDP all-gathers bf16 weights twice and reduce-scatters fp32 gradients, which comes to the same total here:

Model, 8 GPUsTraffic per GPU per stepAt PCIe Gen4 x16, 31.5 GB/s each wayAt A100 NVLink, 300 GB/s each way
Qwen3-1.7B14.2 GB0.45 s0.05 s
Qwen3-8B57.3 GB1.82 s0.19 s

These are floors from our calculation: the traffic from the formulas NVIDIA's nccl-tests use, divided by the link speeds in NVIDIA's A100 whitepaper, 31.5 GB/s per direction for a PCIe Gen4 x16 link and twelve NVLinks of 25 GB/s per direction each. Real transfers are slower, and both DDP and FSDP overlap much of the traffic with computation. The lesson holds anyway: on the L40S and RTX 6000 Ada, whose NVIDIA spec sheets list no NVLink, communication can take longer than the math, so fewer, bigger optimizer steps pay off. With DDP, accumulate gradients inside no_sync(), which skips the all-reduce until the last micro-batch. FSDP's set_requires_gradient_sync(False) skips the reduce-scatter the same way, but it still gathers the weights for every micro-batch, and until the sync each GPU holds the full, unsharded gradients in the reduce dtype: 32.8 GB for Qwen3-8B in fp32 (our calculation: 8.19 billion x 4 bytes), more than a 48 GB card has left next to its 16.4 GB of model states.

Measure the link rather than trusting the floor. This all-reduce test needs nothing beyond the PyTorch environment and uses the bus bandwidth formula from NVIDIA's nccl-tests:

Python
# allreduce_bench.py: torchrun --standalone --nproc-per-node=gpu allreduce_bench.py
import os
import time

import torch
import torch.distributed as dist

local_rank = int(os.environ["LOCAL_RANK"])
torch.cuda.set_device(local_rank)
dist.init_process_group("nccl", device_id=torch.device("cuda", local_rank))
n = dist.get_world_size()
x = torch.zeros(512 * 2**20, dtype=torch.bfloat16, device="cuda")  # 1 GiB
for _ in range(5):
    dist.all_reduce(x)
torch.cuda.synchronize()
t = time.perf_counter()
for _ in range(20):
    dist.all_reduce(x)
torch.cuda.synchronize()
seconds = (time.perf_counter() - t) / 20
busbw = x.numel() * x.element_size() / seconds * 2 * (n - 1) / n / 1e9
if dist.get_rank() == 0:
    print(f"all_reduce bus bandwidth: {busbw:.1f} GB/s")
dist.destroy_process_group()

NVIDIA's own nccl-tests report the same figure over a sweep of message sizes with ./build/all_reduce_perf -b 8 -e 4G -f 2 -g 8, but building them needs the CUDA toolkit and NCCL's development headers on the VM, and the SXM vs PCIe guide walks through the build. If a run is slow or hangs, NCCL_DEBUG=INFO prints NCCL's setup, including the transport it chose between GPUs. NCCL_P2P_DISABLE=1 forces the shared-memory path, which is a test to compare against, not a setting to keep.

The InfiniBand vs NVLink guide explains which link carries which traffic, inside one server and between servers.

Get checkpoints off the VM#

Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots, so a checkpoint only counts once it is somewhere else. The script writes two kinds of output:

OutputContentsSize for Qwen3-1.7BSize for Qwen3-8B
step_50, step_100, ...DCP shards, one file per rank: fp32 weights and both AdamW momentsAbout 25.6 GBAbout 98.3 GB
finalA Hugging Face model folder in bf164.1 GB16.4 GB

The sizes are our calculation: 12 bytes per parameter for the checkpoints, plus, for Qwen3-1.7B, a second fp32 copy of the 311 million embedding weights it shares with its output layer, because DCP saves them under both names. final is the size of the BF16 checkpoint on Hugging Face. A DCP checkpoint loads into any number of GPUs, because DCP reshards at load time, so a run saved on eight L40S can resume on four RTX PRO 6000, with a different data order, since the script splits the data by GPU count. Push the output to a private Hugging Face repo from a second terminal while training runs, every 30 minutes:

Terminal
hf auth login
hf upload your-username/qwen3-run ./ck-fsdp-8b . --private --every=30

At 98 GB per checkpoint for an 8B model, every upload takes a while, so save less often on long runs with --save-every, or copy checkpoints to your own storage with rsync or rclone (moving files to and from a GPU server). To resume on a new VM, download the checkpoint folder into the same --ckpt-dir and run the same torchrun command: the script loads the newest step_ folder. python -m torch.distributed.checkpoint.format_utils dcp_to_torch <checkpoint folder> <file.pt> turns a DCP folder into a single torch.save file when another tool needs one.

Two billing rules protect a long run. The first hour is charged at launch, each further hour when the previous one is used up, and unused seconds are refunded when you stop. If the balance cannot cover the next hour, the instance is terminated and its disk deleted, so fund the run or turn on auto top-up first (pricing).

FAQ#

Should I use DataParallel instead?

No. PyTorch's DDP tutorial says DataParallel is usually slower than DistributedDataParallel even on a single machine, because it runs one process with threads and copies the model on every iteration. DDP with one process per GPU is the standard.

No, it runs over PCIe, but with weights and gradients in the same precision it moves 1.5 times the data DDP does, per the ZeRO paper, so it feels a slow link more. On a PCIe-only VM, use DDP whenever the model fits, raise the tokens per optimizer step with gradient accumulation, and use FSDP only for what DDP cannot hold.

Can I use Hugging Face Trainer or Accelerate instead of a raw loop?

Yes. Both run on the same torchrun launch and support DDP and FSDP2 (Accelerate's fsdp_version: 2), and TRL uses Accelerate underneath. The DeepSpeed ZeRO guide shows that path with TRL, and Axolotl wraps it in one YAML file.

What about more than eight GPUs?

On-demand VMs go up to eight GPUs in one machine, and multi-node training is not self-serve. A full fine-tune of a 70B model needs about 1.1 to 1.3 TB of model states, which is a cluster. We order and build GPU clusters to your spec, with InfiniBand between nodes, and quote the configuration, lead time and terms in writing: send a capacity brief.


My rule for one multi-GPU VM: if the model, its gradients and AdamW state fit one GPU with room for activations, run DDP. If they do not, run FSDP across all the GPUs in the VM, and rent the NVLink-flagged 8x A100 SXM4 when the model is big enough that the traffic dominates. Check nvidia-smi topo -m first, measure the all-reduce, and push every checkpoint off the VM before you stop it. The fine-tuning overview links the other ways to train.

Launch 8x A100 SXM4

Keep building

Choose your next step.