GPU guide

DeepSpeed ZeRO on one multi-GPU VM: stages, memory and configs

DeepSpeed ZeRO stages and offload, memory per GPU from the ZeRO paper and DeepSpeed's estimator, configs, launching on 2 to 8 GPUs, and when to use FSDP.

Faiz Ahmed13 min read

DeepSpeed ZeRO splits training state across GPUs in three stages: stage 1 shards the optimizer state, stage 2 also shards the gradients, and stage 3 also shards the weights. Offload moves the optimizer state, and in stage 3 the weights too, into CPU memory. For a full fine-tune of Qwen3-8B on one 8-GPU VM, DeepSpeed's own estimator puts the model states at 32.42 GiB per GPU under ZeRO-2 and 19.48 GiB under ZeRO-3. So ZeRO-2 wants the 80 GB A100s, and ZeRO-3 leaves room on eight 48 GB L40S or RTX 6000 Ada cards. Pick the lowest stage that fits: stages 1 and 2 add no communication over plain data parallelism, and stage 3 moves up to 1.5 times as much data, according to the ZeRO paper.

What each ZeRO stage shards#

Each stage removes one more copy of the training state from every GPU. The ZeRO paper counts mixed-precision Adam at 16 bytes per parameter: 2 for 16-bit weights, 2 for 16-bit gradients and 12 for the fp32 master weights and the two Adam moments. With P parameters on N GPUs:

StageSharded across GPUsModel states per GPUQwen3-8B on 8 GPUs (our calculation)Communication vs plain data parallel
0, plain data parallelNothing16 x P131.1 GB1x
1Optimizer state4 x P + 12 x P / N45.0 GB1x
2Optimizer state and gradients2 x P + 14 x P / N30.7 GB1x
3Optimizer state, gradients and weights16 x P / N16.4 GBUp to 1.5x

Qwen3-8B has 8.19 billion parameters. Activations, temporary buffers and fragmentation come on top of every figure here, and they grow with sequence length and batch size. The paper's table is the clean version. The next section is what DeepSpeed itself predicts, which includes the buffers it keeps.

Memory per GPU from DeepSpeed's estimator#

DeepSpeed ships estimator functions that need only the parameter count and the largest layer. For Qwen3-8B the largest layer is the 151,936 x 4,096 embedding, 622 million parameters:

Terminal
python -c 'from deepspeed.runtime.zero.stage3 import estimate_zero3_model_states_mem_needs_all_cold; \
estimate_zero3_model_states_mem_needs_all_cold(total_params=8190735360, largest_layer_params=622329856, num_gpus_per_node=8, num_nodes=1)'

DeepSpeed 0.19.7's two estimators print the following for 2, 4 and 8 GPUs. The tool labels its figures GB, but it divides bytes by 2^30, so they are GiB:

Setup, Qwen3-8B full fine-tunePer GPU, 8 GPUsPer GPU, 4 GPUsPer GPU, 2 GPUsCPU memory, 8 GPUs
ZeRO-232.42 GiB49.58 GiB83.91 GiB366.15 GiB while loading
ZeRO-2, optimizer on CPU15.26 GiB15.26 GiB15.26 GiB366.15 GiB
ZeRO-319.48 GiB36.65 GiB70.97 GiB27.82 GiB
ZeRO-3, optimizer on CPU4.23 GiB6.13 GiB9.95 GiB183.08 GiB
ZeRO-3, optimizer and weights on CPU2.32 GiB2.32 GiB2.32 GiB205.96 GiB

The ZeRO-3 rows assume zero.Init, which loads the model already sharded. Transformers turns it on when the DeepSpeed config exists before the model loads, which is how the script below is written. Here is what that leaves on QuantaCloud's multi-GPU VMs (our calculation, from the memory each GPU reports in nvidia-smi, with ECC on for the 48 GB cards):

VMMemory per GPU in nvidia-smiZeRO-2, room leftZeRO-3, room leftPick
8x L40S or 8x RTX 6000 Ada46,068 MiB = 45.0 GiB12.6 GiB25.5 GiBZeRO-3
8x A100 SXM4 80GB81,920 MiB = 80.0 GiB47.6 GiB60.5 GiBZeRO-2
4x RTX PRO 6000 Blackwell97,887 MiB = 95.6 GiB46.0 GiB58.9 GiBZeRO-2
2x H200 NVL143,771 MiB = 140.4 GiB56.5 GiB69.4 GiBZeRO-2

The room left has to hold activations, the logits over a 151,936-token vocabulary and DeepSpeed's communication buckets, so 12 GiB on a 48 GB card is too thin for my taste, and 25 GiB is the margin I would want at 2,048 tokens with gradient checkpointing. The VRAM guide has the same arithmetic for LoRA and QLoRA, where a frozen base model rarely needs ZeRO at all.

GPUMemoryFromAvailable now
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
RTX PRO 6000 Blackwell-Not listedNo
H200 NVL-Not listedNo

Prices checked 5 Oct 2026, 21:12 UTC

Set up DeepSpeed and TRL#

DeepSpeed installs from PyPI as source and compiles its C++ and CUDA extensions the first time a feature needs them. Install PyTorch first, matched to the VM's driver, then the rest:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/ds-env --python 3.12
source ~/ds-env/bin/activate
uv pip install torch --torch-backend=auto
uv pip install deepspeed trl datasets
ds_report

ds_report lists DeepSpeed's ops and, in the environment info at the end, an nvcc version line. The ZeRO-2 and ZeRO-3 configs below name no optimizer, so the Hugging Face Trainer builds PyTorch's AdamW and DeepSpeed has no optimizer kernel to compile. The offload config is the exception: with the optimizer state in CPU memory, the update runs on the CPU in DeepSpeed's own CPU Adam. That config names the optimizer, so DeepSpeed builds CPU Adam directly, and Accelerate, which the Trainer runs on, would swap PyTorch's AdamW for it anyway. CPU Adam compiles on first use with the VM's C++ compiler and, on a GPU machine, links against the CUDA toolkit's libraries. ds_report marks cpu_adam as compatible either way, so read the nvcc version line instead: if it says FAIL, the VM has no CUDA toolkit, and you install one (the driver and CUDA guide shows how) or run inside a CUDA devel container before you use the offload config. The current releases are DeepSpeed 0.19.7, TRL 1.14.0, Transformers 5.17.0 and Accelerate 1.15.0.

The configs#

These start from the Transformers docs' ZeRO configs. "auto" lets the Trainer fill in the batch size, accumulation steps, clipping and precision from its own arguments, so the two never disagree. Save them next to the script.

ds_zero2.json:

json
{
  "bf16": { "enabled": "auto" },
  "zero_optimization": {
    "stage": 2,
    "overlap_comm": true,
    "allgather_bucket_size": 2e8,
    "reduce_bucket_size": 2e8,
    "contiguous_gradients": true
  },
  "gradient_clipping": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "train_batch_size": "auto",
  "gradient_accumulation_steps": "auto"
}

ds_zero3.json:

json
{
  "bf16": { "enabled": "auto" },
  "zero_optimization": {
    "stage": 3,
    "overlap_comm": true,
    "contiguous_gradients": true,
    "reduce_bucket_size": "auto",
    "stage3_prefetch_bucket_size": "auto",
    "stage3_param_persistence_threshold": "auto",
    "stage3_gather_16bit_weights_on_model_save": true
  },
  "gradient_clipping": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "train_batch_size": "auto",
  "gradient_accumulation_steps": "auto"
}

ds_zero3_offload.json adds an optimizer block, filled from the Trainer's arguments, and moves the optimizer state and its update to the CPU:

json
{
  "bf16": { "enabled": "auto" },
  "optimizer": {
    "type": "AdamW",
    "params": { "lr": "auto", "betas": "auto", "eps": "auto", "weight_decay": "auto" }
  },
  "zero_optimization": {
    "stage": 3,
    "overlap_comm": true,
    "contiguous_gradients": true,
    "reduce_bucket_size": "auto",
    "stage3_prefetch_bucket_size": "auto",
    "stage3_param_persistence_threshold": "auto",
    "stage3_gather_16bit_weights_on_model_save": true,
    "offload_optimizer": { "device": "cpu", "pin_memory": true }
  },
  "gradient_clipping": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "train_batch_size": "auto",
  "gradient_accumulation_steps": "auto"
}

Add "offload_param": { "device": "cpu", "pin_memory": true } next to offload_optimizer to move the weights to CPU memory too, which only stage 3 allows. For ZeRO-1, take ds_zero2.json and change the stage to 1. Pinned memory is faster to copy, but DeepSpeed's docs note it counts against the ulimit -l memlock limit, so set "pin_memory": false if the job fails at start-up with an out-of-memory error on the host.

The training script#

TRL's SFTTrainer runs on top of the Hugging Face Trainer, which hands the model to DeepSpeed. This script full fine-tunes Qwen3-8B on TRL's Capybara conversations. Save it as sft_ds.py:

Python
# sft_ds.py: full fine-tune of Qwen3-8B with TRL and DeepSpeed ZeRO
import argparse

import torch
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

parser = argparse.ArgumentParser()
parser.add_argument("--ds-config", default="ds_zero3.json")
parser.add_argument("--model", default="Qwen/Qwen3-8B")
parser.add_argument("--max-steps", type=int, default=-1)  # -1 means one full epoch
args, _ = parser.parse_known_args()  # the deepspeed launcher also passes --local_rank

config = SFTConfig(
    output_dir="qwen3-8b-sft",
    deepspeed=args.ds_config,  # set before the model loads, so ZeRO-3 shards it while loading
    per_device_train_batch_size=1,
    gradient_accumulation_steps=4,
    max_length=2048,
    num_train_epochs=1,
    max_steps=args.max_steps,
    learning_rate=1e-5,
    bf16=True,
    gradient_checkpointing=True,
    logging_steps=10,
    save_strategy="no" if args.max_steps > 0 else "steps",  # a timing run writes no checkpoint
    save_steps=200,
    save_total_limit=2,
    report_to="none",
)
trainer = SFTTrainer(
    model=args.model,  # a model name, so TRL loads it after the DeepSpeed config exists
    args=config,
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
result = trainer.train()
trainer.save_model("qwen3-8b-sft/final")  # ZeRO-3 gathers the 16-bit weights for this
if trainer.is_world_process_zero():
    tokens = [e["num_tokens"] for e in trainer.state.log_history if "num_tokens" in e][-1]
    print(f"tokens_per_s = {tokens / result.metrics['train_runtime']:.0f}")
    print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved() / 2**30:.1f}")

The batch on eight GPUs is 1 per GPU x 4 accumulation steps x 8 GPUs = 32 conversations per optimizer step, which DeepSpeed's config docs define as train_batch_size. One epoch of Capybara's 15,806 conversations is 494 steps (our calculation: 15,806 / 32).

Launch it on one VM#

The DeepSpeed launcher and torchrun both start one process per GPU, and the Transformers docs list both. They also list Accelerate's launcher, but it ignores the deepspeed setting in the script's config and needs a config file of its own, so this guide uses the first two:

Terminal
# ZeRO-3 on every GPU in the VM
deepspeed --num_gpus 8 sft_ds.py --ds-config ds_zero3.json
# the same job with torchrun
torchrun --standalone --nproc-per-node=gpu sft_ds.py --ds-config ds_zero3.json
# ZeRO-2 on only GPUs 0 to 3
deepspeed --include localhost:0,1,2,3 sft_ds.py --ds-config ds_zero2.json

Start with --max-steps 50 to check memory and speed before you commit to the full epoch. A run with --max-steps writes no checkpoints, because the Trainer's timer would otherwise include saving about 115 GB of DeepSpeed state at the end (our calculation: 16.4 GB of bf16 weights plus 98.3 GB of fp32 optimizer state). The script prints tokens_per_s and rank 0's peak memory at the end, the Trainer logs the loss every 10 steps, and the cost guide turns tokens per second into hours and dollars. The DDP and FSDP guide shows how to check the links between GPUs with nvidia-smi topo -m and an all-reduce test, which matters most for ZeRO-3 on PCIe-only VMs such as the L40S.

Launch 8x A100 SXM4 for ZeRO-2 Launch 8x L40S for ZeRO-3

Checkpoints and getting the weights out#

Every save_steps, and once more when training ends, the Trainer writes a checkpoint-N folder with DeepSpeed's sharded state, one set of files per rank. The optimizer part alone is 12 bytes per parameter, about 98 GB for Qwen3-8B (our calculation), which is why save_total_limit=2 keeps only the newest two. DeepSpeed also drops a zero_to_fp32.py script into each checkpoint, which rebuilds full fp32 weights without GPUs:

Terminal
cd qwen3-8b-sft/checkpoint-400
python zero_to_fp32.py . ../fp32-weights --safe_serialization

For deployment you rarely need that, because stage3_gather_16bit_weights_on_model_save makes trainer.save_model write a normal 16-bit Hugging Face model to qwen3-8b-sft/final. DeepSpeed's docs warn that gathering the weights onto one GPU is slow and memory hungry, so do it once at the end.

Stopping a QuantaCloud instance deletes its disk, and nothing on it survives. Push the output to a private Hugging Face repo from a second terminal, every 30 minutes, while training runs:

Terminal
hf auth login
hf upload your-username/qwen3-8b-sft ./qwen3-8b-sft . --private --every=30

The file transfer guide covers rsync and rclone to your own storage instead. Before a long run, fund it or turn on auto top-up: the first hour is charged at launch, each further hour when the previous one is used up, and if the balance cannot cover the next hour the instance is terminated with its disk (pricing).

When to use FSDP instead#

I would use FSDP2 for new code and keep DeepSpeed for the jobs where its offload earns its keep. The two overlap more than their names suggest: Accelerate's comparison maps FSDP's full sharding to ZeRO stage 3, and both run on the same VM with the same launchers. The differences that decide it:

QuestionDeepSpeed ZeROFSDP2 (fully_shard)
Extra softwareA separate package whose ops compile on first usePart of PyTorch
OffloadOptimizer and weights separately, to CPU or NVMeAll or nothing, to CPU
Master weightsfp32 by default, and Accelerate warns the upcast can cost memory on few GPUsThe dtype the model is loaded in, fp32 in the DDP and FSDP guide's script
Saving a 16-bit modelOne config flag gathers it at save timeGather a full state dict with PyTorch's checkpoint APIs
Tool supportTransformers, TRL, Accelerate, AxolotlThe same tools. Axolotl recommends FSDP2 for new users

Use DeepSpeed when a model only fits with optimizer offload, when you want NVMe offload, or when the recipe you are following already ships a DeepSpeed config. Use FSDP2 when the model fits sharded across the VM's GPUs and you want one less dependency. For LoRA and QLoRA on one VM, plain DDP is often enough, and Axolotl switches between DDP, DeepSpeed and FSDP with one line of YAML.

FAQ#

What is DeepSpeed?

DeepSpeed is an open-source library for training large models on many GPUs. Its best-known part is ZeRO, the Zero Redundancy Optimizer, which removes the duplicate copies of training state that plain data parallelism keeps on every GPU. It plugs into the Hugging Face Trainer through a JSON config, as above.

Which ZeRO stage should I start with?

The lowest one that fits. Axolotl's docs say to go from stage 1 to 2 to 3, and the Transformers docs say to use ZeRO-3 only when the model does not fit with ZeRO-2. Stages 1 and 2 cost no extra communication, and stage 3 costs up to 1.5 times as much.

No, but it moves the most data of any stage, so it gains the most from a fast link. On a PCIe-only VM such as 8x L40S, ZeRO-3 still runs, and the extra traffic is the price of fitting the model. Check the links with nvidia-smi topo -m, and read SXM vs PCIe for what each link between GPUs carries.

Can DeepSpeed train across several VMs?

The launcher supports multiple nodes with a hostfile, but on-demand QuantaCloud VMs are single machines of up to eight GPUs, and multi-node training is not self-serve. For jobs that need more, we order and build GPU clusters to your spec, with InfiniBand between nodes, and quote the configuration, lead time and terms in writing: send a capacity brief.


The rule I follow for ZeRO: run the estimator for your model and GPU count, take the lowest stage that leaves room for activations, and add offload only when nothing else fits. For Qwen3-8B that means ZeRO-2 on an 8x A100 SXM4 and ZeRO-3 on an 8x L40S. If the model fits sharded without offload, FSDP2 does the same job inside PyTorch, and the fine-tuning overview links every other way to train.

Launch 8x L40S

Keep building

Choose your next step.