Use DDP when one GPU can hold the whole model with its gradients and optimizer state, and FSDP when it cannot. Both launch the same way on one multi-GPU VM, one process per GPU with torchrun --standalone --nproc-per-node=gpu, and the script below runs either with one flag. The difference is memory: a full fine-tune of Qwen3-1.7B holds 32.5 GB of weights, gradients and AdamW state on every GPU under DDP, and 4.1 GB per GPU under FSDP across eight (our calculation: 2.03 billion parameters x 16 bytes, then divided by 8). The price FSDP pays is communication: it gathers each layer's weights from the other GPUs before using them, on top of reducing the gradients, and that traffic is where NVLink, or its absence, starts to matter.
DDP or FSDP: the rule#
The rule I follow is to start with DDP and move to FSDP only when DDP runs out of memory, because DDP is simpler and moves no more data per step. The two differ in what each GPU keeps and what crosses between GPUs:
DDP (DistributedDataParallel) | FSDP2 (fully_shard) | |
|---|---|---|
| Each GPU holds | A full copy of the weights, gradients and optimizer state | 1/N of each, plus the layer it is computing |
| Communication per step | One all-reduce of the gradients during backward | An all-gather of each layer's weights in forward and again in backward, and a reduce-scatter of its gradients |
| Model states per GPU, AdamW with fp32 master weights | 16 bytes per parameter, 20 with DDP's default gradient buckets | 16 bytes per parameter divided by N |
| Largest full fine-tune, model states only | About 3B parameters on a 48 GB GPU and 5B on 80 GB (our calculation: memory / 16 bytes) | Up to the VM's total GPU memory / 16 bytes |
| Code change | Wrap the model in DDP(...) | Call fully_shard on each layer, then on the whole model |
The 16 bytes are the ZeRO paper's accounting for mixed-precision Adam, and FSDP is the same idea as its stage 3. PyTorch's own tutorial describes FSDP as DDP's all-reduce split into a reduce-scatter and an all-gather, and the ZeRO paper puts full sharding at 1.5 times the communication volume of plain data parallelism. FSDP1, the old FullyShardedDataParallel class, is deprecated. Use FSDP2's fully_shard, as below.
Memory per GPU for a full fine-tune#
The table is our calculation of model states only: weights, gradients and AdamW state at 16 bytes per parameter. Activations come on top and grow with sequence length and batch size, which is why the script recomputes them in the backward pass.
| Model | DDP, every GPU | FSDP on 2 GPUs | FSDP on 4 GPUs | FSDP on 8 GPUs |
|---|---|---|---|---|
| Qwen3-1.7B (2.03B parameters) | 32.5 GB | 16.3 GB | 8.1 GB | 4.1 GB |
| Qwen3-8B (8.19B parameters) | 131.1 GB | 65.5 GB | 32.8 GB | 16.4 GB |
Two readings matter. Qwen3-1.7B fits DDP on a 48 GB GPU with 15.5 GB left for activations (our calculation: 48 - 32.5), as long as you pass gradient_as_bucket_view=True: PyTorch's docs say it saves memory equal to the total gradient size, which is another 8.1 GB here. Qwen3-8B does not fit DDP on anything short of the 141 GB H200 NVL, and even there it leaves about 18 GiB for activations, so it is an FSDP job: 16.4 GB per GPU on eight 48 GB L40S or RTX 6000 Ada cards (our calculation: 140.4 GiB reported minus 131.1 GB, 122.1 GiB, and 131.1 / 8). For LoRA and QLoRA, where the base model is frozen, the VRAM guide has the numbers, and most of those jobs need one GPU and plain DDP at most.
Pick the multi-GPU VM#
Any QuantaCloud GPU instance with 2, 4 or 8 GPUs runs this. On 2026-09-27 the largest configurations were:
| Configuration | GPU memory | RAM | Link between GPUs |
|---|---|---|---|
| 8x A100 SXM4 80GB | 640 GB | 800 GB | SXM, flagged NVLink by the API |
| 8x L40S | 384 GB | 576 GB | PCIe Gen4 x16, no NVLink on this card |
| 8x RTX 6000 Ada | 384 GB | 640 GB | PCIe 4.0 x16, no NVLink on this card |
| 4x RTX PRO 6000 Blackwell | 384 GB | 576 GB | PCIe, no NVLink on this card |
| 2x H200 NVL | 282 GB | 360 GB | Flagged NVLink by the API |
RAM matters because every process loads the full model into CPU memory before sharding it: eight fp32 copies of Qwen3-8B are 262 GB (our calculation: 8 x 8.19 billion x 4 bytes). Live prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| A100 SXM4 80GB | 80 GB | $1.50/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 13:48 UTC
Launch 8x A100 SXM4 Launch 8x L40SEight-GPU VMs take about 10 minutes to reach running (median). Connect as ubuntu over SSH (SSH docs), or with VS Code Remote. The A100 page and the L40S page cover the cards themselves.
Check how the GPUs connect#
The first thing I check on a multi-GPU VM is the link matrix, because it decides how fast DDP and FSDP can exchange data:
nvidia-smi
nvidia-smi topo -m
nvidia-smi topo -p2p p
In topo -m, an entry of NV plus a number between two GPUs means that many bonded NVLinks. PIX, PXB, PHB, NODE and SYS mean PCIe, with the path getting longer in that order, and SYS crosses between NUMA nodes. topo -p2p p prints OK for each pair of GPUs that supports peer-to-peer transfers over PCIe. When that peer-to-peer path is missing, NCCL falls back to its shared-memory transport, which copies through host memory. NCCL's troubleshooting guide notes two things that apply to any VM: virtual machines need PCI Access Control Services turned on, so the usual bare-metal fix of disabling it is not available, and a virtual PCI topology can cost performance. QuantaCloud's API flags the A100 SXM4 and H200 NVL offers as NVLink, and the matrix is how you confirm it on your own VM.
Set up the environment#
A virtual environment with PyTorch, Transformers and datasets is all the script needs. uv picks the PyTorch CUDA build that matches the VM's driver:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/dist-env --python 3.12
source ~/dist-env/bin/activate
uv pip install torch --torch-backend=auto
uv pip install transformers datasets
python -c "import torch; print(torch.__version__, torch.cuda.device_count(), torch.cuda.nccl.version())"
The last line should print the PyTorch version, the number of GPUs in the VM and the NCCL version bundled with PyTorch. The current releases are PyTorch 2.14.0 and Transformers 5.17.0. The driver and CUDA guide explains why the driver decides which PyTorch build you get.
The training script#
The script below full fine-tunes a Hugging Face model on TRL's Capybara conversations, with --mode ddp or --mode fsdp. It keeps fp32 master weights and computes in bf16, recomputes activations in the backward pass, saves a sharded checkpoint every 50 steps, resumes from the newest one, and writes a normal Hugging Face model folder at the end. Save it as train.py:
# train.py: full fine-tuning of a Hugging Face model with DDP or FSDP2 on one VM
# Launch: torchrun --standalone --nproc-per-node=gpu train.py --mode fsdp
import argparse
import contextlib
import os
import time
import torch
import torch.distributed as dist
import torch.distributed.checkpoint as dcp
from datasets import load_dataset
from torch.distributed.checkpoint.state_dict import (
StateDictOptions, get_model_state_dict, get_state_dict, set_state_dict)
from torch.distributed.checkpoint.stateful import Stateful
from torch.distributed.fsdp import MixedPrecisionPolicy, fully_shard
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.utils.data import DataLoader, DistributedSampler
from transformers import AutoModelForCausalLM, AutoTokenizer
parser = argparse.ArgumentParser()
parser.add_argument("--mode", choices=["ddp", "fsdp"], default="fsdp")
parser.add_argument("--model", default="Qwen/Qwen3-1.7B")
parser.add_argument("--seq-len", type=int, default=2048)
parser.add_argument("--steps", type=int, default=100)
parser.add_argument("--lr", type=float, default=1e-5)
parser.add_argument("--save-every", type=int, default=50)
parser.add_argument("--ckpt-dir", default="checkpoints")
args = parser.parse_args()
# torchrun starts one process per GPU and sets LOCAL_RANK, RANK and WORLD_SIZE
local_rank = int(os.environ["LOCAL_RANK"])
device = torch.device("cuda", local_rank)
torch.cuda.set_device(device)
dist.init_process_group(backend="nccl", device_id=device)
rank, world = dist.get_rank(), dist.get_world_size()
# Data: TRL's Capybara conversations, tokenized and packed into seq-len blocks
tok = AutoTokenizer.from_pretrained(args.model)
texts = [tok.apply_chat_template(m, tokenize=False)
for m in load_dataset("trl-lib/Capybara", split="train")["messages"]]
flat = [t for ids in tok(texts, add_special_tokens=False)["input_ids"] for t in ids]
n = len(flat) // args.seq_len
blocks = torch.tensor(flat[: n * args.seq_len]).view(n, args.seq_len)
sampler = DistributedSampler(blocks, shuffle=True, seed=0) # a different slice per rank
loader = DataLoader(blocks, batch_size=1, sampler=sampler, drop_last=True)
# Model: fp32 master weights, bf16 compute, activations recomputed in backward
model = AutoModelForCausalLM.from_pretrained(args.model, dtype=torch.float32, use_cache=False)
model.gradient_checkpointing_enable()
if args.mode == "ddp":
# every GPU holds the whole model, its gradients and the optimizer state
model = DDP(model.to(device), device_ids=[local_rank], gradient_as_bucket_view=True)
amp = lambda: torch.autocast("cuda", dtype=torch.bfloat16)
else:
# each GPU holds 1/N of them; layers are all-gathered just in time
mp = MixedPrecisionPolicy(param_dtype=torch.bfloat16, reduce_dtype=torch.float32)
for layer in model.model.layers:
fully_shard(layer, mp_policy=mp)
fully_shard(model, mp_policy=mp)
amp = contextlib.nullcontext
optimizer = torch.optim.AdamW(model.parameters(), lr=args.lr, weight_decay=0.0)
class AppState(Stateful): # lets DCP save and load model and optimizer shards
def __init__(self, model, optimizer):
self.model, self.optimizer = model, optimizer
def state_dict(self):
model_sd, optim_sd = get_state_dict(self.model, self.optimizer)
return {"model": model_sd, "optim": optim_sd}
def load_state_dict(self, sd):
set_state_dict(self.model, self.optimizer,
model_state_dict=sd["model"], optim_state_dict=sd["optim"])
# Resume from the newest checkpoint folder if one exists
step = 0
if os.path.isdir(args.ckpt_dir):
done = sorted(int(d.split("_")[1]) for d in os.listdir(args.ckpt_dir) if d.startswith("step_"))
if done:
step = done[-1]
dcp.load({"app": AppState(model, optimizer)}, checkpoint_id=f"{args.ckpt_dir}/step_{step}")
def batches():
epoch = 0
while True:
sampler.set_epoch(epoch)
yield from loader
epoch += 1
data = batches()
for _ in range(step): # skip the batches the checkpoint already trained on
next(data)
model.train()
torch.cuda.reset_peak_memory_stats(device)
timed_s, timed_tokens = 0.0, 0
while step < args.steps:
t = time.perf_counter()
x = next(data).to(device)
with amp():
loss = model(input_ids=x, labels=x).loss
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
optimizer.zero_grad()
torch.cuda.synchronize(device)
step += 1
if step > 10: # the first steps include warm-up, so they are not timed
timed_s += time.perf_counter() - t
timed_tokens += x.numel() * world
if rank == 0 and step % 10 == 0:
print(f"step {step} loss {loss.item():.3f}", flush=True)
if step % args.save_every == 0 or step == args.steps:
dcp.save({"app": AppState(model, optimizer)}, checkpoint_id=f"{args.ckpt_dir}/step_{step}")
if rank == 0 and timed_s > 0:
print(f"tokens_per_s = {timed_tokens / timed_s:.0f}")
print(f"peak_allocated_gib = {torch.cuda.max_memory_allocated(device) / 2**30:.1f}")
print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved(device) / 2**30:.1f}")
# Final model in Hugging Face format, bf16, written by rank 0
full_sd = get_model_state_dict(model, options=StateDictOptions(full_state_dict=True, cpu_offload=True))
if rank == 0:
final = AutoModelForCausalLM.from_pretrained(args.model, dtype=torch.bfloat16)
final.load_state_dict(full_sd)
final.save_pretrained(f"{args.ckpt_dir}/final")
tok.save_pretrained(f"{args.ckpt_dir}/final")
dist.destroy_process_group()
The lines that decide memory and speed:
| Line | What it does |
|---|---|
DistributedSampler(...) | Gives each process its own share of the data. Without it, every GPU trains on the same batches |
DDP(..., gradient_as_bucket_view=True) | Makes the gradients views into DDP's communication buckets instead of a second copy |
fully_shard(layer, ...) then fully_shard(model, ...) | Shards layer by layer, bottom-up, as PyTorch's docs require, so each layer's full weights are gathered just before it runs and freed after |
MixedPrecisionPolicy(param_dtype=bf16, reduce_dtype=fp32) | Computes in bf16 and reduces gradients in fp32, while the optimizer updates fp32 shards |
AdamW(model.parameters()) after sharding | The optimizer must see the sharded parameters, so it is built after fully_shard |
gradient_checkpointing_enable() | Recomputes activations in backward. Transformers 5 uses the non-reentrant mode by default |
dcp.save(...) | Every rank writes its own shard in parallel, so saving does not funnel through one GPU |
Each step trains on one 2,048-token block per GPU, 16,384 tokens per step on eight GPUs (our calculation: 8 x 2,048). The script trains on whole packed conversations, which keeps it short. For chat fine-tuning with the prompts masked out, Axolotl trains on the assistant turns by default, and TRL, which the DeepSpeed ZeRO guide uses, does it with assistant_only_loss=True.
Launch it with torchrun#
The same script covers every case. Start with one GPU to get a baseline, then use all of them:
# baseline: one GPU
CUDA_VISIBLE_DEVICES=0 torchrun --standalone --nproc-per-node=1 train.py --mode ddp --ckpt-dir ck-1gpu
# DDP on every GPU in the VM
torchrun --standalone --nproc-per-node=gpu train.py --mode ddp --ckpt-dir ck-ddp
# FSDP on every GPU, and a model DDP cannot hold
torchrun --standalone --nproc-per-node=gpu train.py --mode fsdp --model Qwen/Qwen3-8B --ckpt-dir ck-fsdp-8b
--standalone runs the rendezvous on the VM itself, and --nproc-per-node=gpu starts one process per visible GPU. Each run prints its loss every 10 steps and, at the end, tokens_per_s and the peak memory on rank 0. Scaling efficiency is the 8-GPU tokens per second divided by eight times the one-GPU figure. If it comes out low, low GPU utilization covers why GPUs wait on each other, and the cost guide turns tokens per second into hours and dollars. For long runs, start the job with nohup ... > train.log 2>&1 & so a dropped SSH session does not kill it, and watch the GPUs from a second terminal with nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv -l 5.
Communication without NVLink#
The core problem on a PCIe-only VM is that every step moves the same bytes whether the link is fast or not. With this script's settings, DDP all-reduces fp32 gradients, and each GPU sends and receives 2 x (N-1)/N times their size. FSDP all-gathers bf16 weights twice and reduce-scatters fp32 gradients, which comes to the same total here:
| Model, 8 GPUs | Traffic per GPU per step | At PCIe Gen4 x16, 31.5 GB/s each way | At A100 NVLink, 300 GB/s each way |
|---|---|---|---|
| Qwen3-1.7B | 14.2 GB | 0.45 s | 0.05 s |
| Qwen3-8B | 57.3 GB | 1.82 s | 0.19 s |
These are floors from our calculation: the traffic from the formulas NVIDIA's nccl-tests use, divided by the link speeds in NVIDIA's A100 whitepaper, 31.5 GB/s per direction for a PCIe Gen4 x16 link and twelve NVLinks of 25 GB/s per direction each. Real transfers are slower, and both DDP and FSDP overlap much of the traffic with computation. The lesson holds anyway: on the L40S and RTX 6000 Ada, whose NVIDIA spec sheets list no NVLink, communication can take longer than the math, so fewer, bigger optimizer steps pay off. With DDP, accumulate gradients inside no_sync(), which skips the all-reduce until the last micro-batch. FSDP's set_requires_gradient_sync(False) skips the reduce-scatter the same way, but it still gathers the weights for every micro-batch, and until the sync each GPU holds the full, unsharded gradients in the reduce dtype: 32.8 GB for Qwen3-8B in fp32 (our calculation: 8.19 billion x 4 bytes), more than a 48 GB card has left next to its 16.4 GB of model states.
Measure the link rather than trusting the floor. This all-reduce test needs nothing beyond the PyTorch environment and uses the bus bandwidth formula from NVIDIA's nccl-tests:
# allreduce_bench.py: torchrun --standalone --nproc-per-node=gpu allreduce_bench.py
import os
import time
import torch
import torch.distributed as dist
local_rank = int(os.environ["LOCAL_RANK"])
torch.cuda.set_device(local_rank)
dist.init_process_group("nccl", device_id=torch.device("cuda", local_rank))
n = dist.get_world_size()
x = torch.zeros(512 * 2**20, dtype=torch.bfloat16, device="cuda") # 1 GiB
for _ in range(5):
dist.all_reduce(x)
torch.cuda.synchronize()
t = time.perf_counter()
for _ in range(20):
dist.all_reduce(x)
torch.cuda.synchronize()
seconds = (time.perf_counter() - t) / 20
busbw = x.numel() * x.element_size() / seconds * 2 * (n - 1) / n / 1e9
if dist.get_rank() == 0:
print(f"all_reduce bus bandwidth: {busbw:.1f} GB/s")
dist.destroy_process_group()
NVIDIA's own nccl-tests report the same figure over a sweep of message sizes with ./build/all_reduce_perf -b 8 -e 4G -f 2 -g 8, but building them needs the CUDA toolkit and NCCL's development headers on the VM, and the SXM vs PCIe guide walks through the build. If a run is slow or hangs, NCCL_DEBUG=INFO prints NCCL's setup, including the transport it chose between GPUs. NCCL_P2P_DISABLE=1 forces the shared-memory path, which is a test to compare against, not a setting to keep.
The InfiniBand vs NVLink guide explains which link carries which traffic, inside one server and between servers.
Get checkpoints off the VM#
Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots, so a checkpoint only counts once it is somewhere else. The script writes two kinds of output:
| Output | Contents | Size for Qwen3-1.7B | Size for Qwen3-8B |
|---|---|---|---|
step_50, step_100, ... | DCP shards, one file per rank: fp32 weights and both AdamW moments | About 25.6 GB | About 98.3 GB |
final | A Hugging Face model folder in bf16 | 4.1 GB | 16.4 GB |
The sizes are our calculation: 12 bytes per parameter for the checkpoints, plus, for Qwen3-1.7B, a second fp32 copy of the 311 million embedding weights it shares with its output layer, because DCP saves them under both names. final is the size of the BF16 checkpoint on Hugging Face. A DCP checkpoint loads into any number of GPUs, because DCP reshards at load time, so a run saved on eight L40S can resume on four RTX PRO 6000, with a different data order, since the script splits the data by GPU count. Push the output to a private Hugging Face repo from a second terminal while training runs, every 30 minutes:
hf auth login
hf upload your-username/qwen3-run ./ck-fsdp-8b . --private --every=30
At 98 GB per checkpoint for an 8B model, every upload takes a while, so save less often on long runs with --save-every, or copy checkpoints to your own storage with rsync or rclone (moving files to and from a GPU server). To resume on a new VM, download the checkpoint folder into the same --ckpt-dir and run the same torchrun command: the script loads the newest step_ folder. python -m torch.distributed.checkpoint.format_utils dcp_to_torch <checkpoint folder> <file.pt> turns a DCP folder into a single torch.save file when another tool needs one.
Two billing rules protect a long run. The first hour is charged at launch, each further hour when the previous one is used up, and unused seconds are refunded when you stop. If the balance cannot cover the next hour, the instance is terminated and its disk deleted, so fund the run or turn on auto top-up first (pricing).
FAQ#
Should I use DataParallel instead?
No. PyTorch's DDP tutorial says DataParallel is usually slower than DistributedDataParallel even on a single machine, because it runs one process with threads and copies the model on every iteration. DDP with one process per GPU is the standard.
Does FSDP need NVLink?
No, it runs over PCIe, but with weights and gradients in the same precision it moves 1.5 times the data DDP does, per the ZeRO paper, so it feels a slow link more. On a PCIe-only VM, use DDP whenever the model fits, raise the tokens per optimizer step with gradient accumulation, and use FSDP only for what DDP cannot hold.
Can I use Hugging Face Trainer or Accelerate instead of a raw loop?
Yes. Both run on the same torchrun launch and support DDP and FSDP2 (Accelerate's fsdp_version: 2), and TRL uses Accelerate underneath. The DeepSpeed ZeRO guide shows that path with TRL, and Axolotl wraps it in one YAML file.
What about more than eight GPUs?
On-demand VMs go up to eight GPUs in one machine, and multi-node training is not self-serve. A full fine-tune of a 70B model needs about 1.1 to 1.3 TB of model states, which is a cluster. We order and build GPU clusters to your spec, with InfiniBand between nodes, and quote the configuration, lead time and terms in writing: send a capacity brief.
My rule for one multi-GPU VM: if the model, its gradients and AdamW state fit one GPU with room for activations, run DDP. If they do not, run FSDP across all the GPUs in the VM, and rent the NVLink-flagged 8x A100 SXM4 when the model is big enough that the traffic dominates. Check nvidia-smi topo -m first, measure the all-reduce, and push every checkpoint off the VM before you stop it. The fine-tuning overview links the other ways to train.