DeepSpeed ZeRO splits training state across GPUs in three stages: stage 1 shards the optimizer state, stage 2 also shards the gradients, and stage 3 also shards the weights. Offload moves the optimizer state, and in stage 3 the weights too, into CPU memory. For a full fine-tune of Qwen3-8B on one 8-GPU VM, DeepSpeed's own estimator puts the model states at 32.42 GiB per GPU under ZeRO-2 and 19.48 GiB under ZeRO-3. So ZeRO-2 wants the 80 GB A100s, and ZeRO-3 leaves room on eight 48 GB L40S or RTX 6000 Ada cards. Pick the lowest stage that fits: stages 1 and 2 add no communication over plain data parallelism, and stage 3 moves up to 1.5 times as much data, according to the ZeRO paper.
What each ZeRO stage shards#
Each stage removes one more copy of the training state from every GPU. The ZeRO paper counts mixed-precision Adam at 16 bytes per parameter: 2 for 16-bit weights, 2 for 16-bit gradients and 12 for the fp32 master weights and the two Adam moments. With P parameters on N GPUs:
| Stage | Sharded across GPUs | Model states per GPU | Qwen3-8B on 8 GPUs (our calculation) | Communication vs plain data parallel |
|---|---|---|---|---|
| 0, plain data parallel | Nothing | 16 x P | 131.1 GB | 1x |
| 1 | Optimizer state | 4 x P + 12 x P / N | 45.0 GB | 1x |
| 2 | Optimizer state and gradients | 2 x P + 14 x P / N | 30.7 GB | 1x |
| 3 | Optimizer state, gradients and weights | 16 x P / N | 16.4 GB | Up to 1.5x |
Qwen3-8B has 8.19 billion parameters. Activations, temporary buffers and fragmentation come on top of every figure here, and they grow with sequence length and batch size. The paper's table is the clean version. The next section is what DeepSpeed itself predicts, which includes the buffers it keeps.
Memory per GPU from DeepSpeed's estimator#
DeepSpeed ships estimator functions that need only the parameter count and the largest layer. For Qwen3-8B the largest layer is the 151,936 x 4,096 embedding, 622 million parameters:
python -c 'from deepspeed.runtime.zero.stage3 import estimate_zero3_model_states_mem_needs_all_cold; \
estimate_zero3_model_states_mem_needs_all_cold(total_params=8190735360, largest_layer_params=622329856, num_gpus_per_node=8, num_nodes=1)'
DeepSpeed 0.19.7's two estimators print the following for 2, 4 and 8 GPUs. The tool labels its figures GB, but it divides bytes by 2^30, so they are GiB:
| Setup, Qwen3-8B full fine-tune | Per GPU, 8 GPUs | Per GPU, 4 GPUs | Per GPU, 2 GPUs | CPU memory, 8 GPUs |
|---|---|---|---|---|
| ZeRO-2 | 32.42 GiB | 49.58 GiB | 83.91 GiB | 366.15 GiB while loading |
| ZeRO-2, optimizer on CPU | 15.26 GiB | 15.26 GiB | 15.26 GiB | 366.15 GiB |
| ZeRO-3 | 19.48 GiB | 36.65 GiB | 70.97 GiB | 27.82 GiB |
| ZeRO-3, optimizer on CPU | 4.23 GiB | 6.13 GiB | 9.95 GiB | 183.08 GiB |
| ZeRO-3, optimizer and weights on CPU | 2.32 GiB | 2.32 GiB | 2.32 GiB | 205.96 GiB |
The ZeRO-3 rows assume zero.Init, which loads the model already sharded. Transformers turns it on when the DeepSpeed config exists before the model loads, which is how the script below is written. Here is what that leaves on QuantaCloud's multi-GPU VMs (our calculation, from the memory each GPU reports in nvidia-smi, with ECC on for the 48 GB cards):
| VM | Memory per GPU in nvidia-smi | ZeRO-2, room left | ZeRO-3, room left | Pick |
|---|---|---|---|---|
| 8x L40S or 8x RTX 6000 Ada | 46,068 MiB = 45.0 GiB | 12.6 GiB | 25.5 GiB | ZeRO-3 |
| 8x A100 SXM4 80GB | 81,920 MiB = 80.0 GiB | 47.6 GiB | 60.5 GiB | ZeRO-2 |
| 4x RTX PRO 6000 Blackwell | 97,887 MiB = 95.6 GiB | 46.0 GiB | 58.9 GiB | ZeRO-2 |
| 2x H200 NVL | 143,771 MiB = 140.4 GiB | 56.5 GiB | 69.4 GiB | ZeRO-2 |
The room left has to hold activations, the logits over a 151,936-token vocabulary and DeepSpeed's communication buckets, so 12 GiB on a 48 GB card is too thin for my taste, and 25 GiB is the margin I would want at 2,048 tokens with gradient checkpointing. The VRAM guide has the same arithmetic for LoRA and QLoRA, where a frozen base model rarely needs ZeRO at all.
| GPU | Memory | From | Available now |
|---|---|---|---|
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 21:12 UTC
Set up DeepSpeed and TRL#
DeepSpeed installs from PyPI as source and compiles its C++ and CUDA extensions the first time a feature needs them. Install PyTorch first, matched to the VM's driver, then the rest:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/ds-env --python 3.12
source ~/ds-env/bin/activate
uv pip install torch --torch-backend=auto
uv pip install deepspeed trl datasets
ds_report
ds_report lists DeepSpeed's ops and, in the environment info at the end, an nvcc version line. The ZeRO-2 and ZeRO-3 configs below name no optimizer, so the Hugging Face Trainer builds PyTorch's AdamW and DeepSpeed has no optimizer kernel to compile. The offload config is the exception: with the optimizer state in CPU memory, the update runs on the CPU in DeepSpeed's own CPU Adam. That config names the optimizer, so DeepSpeed builds CPU Adam directly, and Accelerate, which the Trainer runs on, would swap PyTorch's AdamW for it anyway. CPU Adam compiles on first use with the VM's C++ compiler and, on a GPU machine, links against the CUDA toolkit's libraries. ds_report marks cpu_adam as compatible either way, so read the nvcc version line instead: if it says FAIL, the VM has no CUDA toolkit, and you install one (the driver and CUDA guide shows how) or run inside a CUDA devel container before you use the offload config. The current releases are DeepSpeed 0.19.7, TRL 1.14.0, Transformers 5.17.0 and Accelerate 1.15.0.
The configs#
These start from the Transformers docs' ZeRO configs. "auto" lets the Trainer fill in the batch size, accumulation steps, clipping and precision from its own arguments, so the two never disagree. Save them next to the script.
ds_zero2.json:
{
"bf16": { "enabled": "auto" },
"zero_optimization": {
"stage": 2,
"overlap_comm": true,
"allgather_bucket_size": 2e8,
"reduce_bucket_size": 2e8,
"contiguous_gradients": true
},
"gradient_clipping": "auto",
"train_micro_batch_size_per_gpu": "auto",
"train_batch_size": "auto",
"gradient_accumulation_steps": "auto"
}
ds_zero3.json:
{
"bf16": { "enabled": "auto" },
"zero_optimization": {
"stage": 3,
"overlap_comm": true,
"contiguous_gradients": true,
"reduce_bucket_size": "auto",
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_gather_16bit_weights_on_model_save": true
},
"gradient_clipping": "auto",
"train_micro_batch_size_per_gpu": "auto",
"train_batch_size": "auto",
"gradient_accumulation_steps": "auto"
}
ds_zero3_offload.json adds an optimizer block, filled from the Trainer's arguments, and moves the optimizer state and its update to the CPU:
{
"bf16": { "enabled": "auto" },
"optimizer": {
"type": "AdamW",
"params": { "lr": "auto", "betas": "auto", "eps": "auto", "weight_decay": "auto" }
},
"zero_optimization": {
"stage": 3,
"overlap_comm": true,
"contiguous_gradients": true,
"reduce_bucket_size": "auto",
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_gather_16bit_weights_on_model_save": true,
"offload_optimizer": { "device": "cpu", "pin_memory": true }
},
"gradient_clipping": "auto",
"train_micro_batch_size_per_gpu": "auto",
"train_batch_size": "auto",
"gradient_accumulation_steps": "auto"
}
Add "offload_param": { "device": "cpu", "pin_memory": true } next to offload_optimizer to move the weights to CPU memory too, which only stage 3 allows. For ZeRO-1, take ds_zero2.json and change the stage to 1. Pinned memory is faster to copy, but DeepSpeed's docs note it counts against the ulimit -l memlock limit, so set "pin_memory": false if the job fails at start-up with an out-of-memory error on the host.
The training script#
TRL's SFTTrainer runs on top of the Hugging Face Trainer, which hands the model to DeepSpeed. This script full fine-tunes Qwen3-8B on TRL's Capybara conversations. Save it as sft_ds.py:
# sft_ds.py: full fine-tune of Qwen3-8B with TRL and DeepSpeed ZeRO
import argparse
import torch
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer
parser = argparse.ArgumentParser()
parser.add_argument("--ds-config", default="ds_zero3.json")
parser.add_argument("--model", default="Qwen/Qwen3-8B")
parser.add_argument("--max-steps", type=int, default=-1) # -1 means one full epoch
args, _ = parser.parse_known_args() # the deepspeed launcher also passes --local_rank
config = SFTConfig(
output_dir="qwen3-8b-sft",
deepspeed=args.ds_config, # set before the model loads, so ZeRO-3 shards it while loading
per_device_train_batch_size=1,
gradient_accumulation_steps=4,
max_length=2048,
num_train_epochs=1,
max_steps=args.max_steps,
learning_rate=1e-5,
bf16=True,
gradient_checkpointing=True,
logging_steps=10,
save_strategy="no" if args.max_steps > 0 else "steps", # a timing run writes no checkpoint
save_steps=200,
save_total_limit=2,
report_to="none",
)
trainer = SFTTrainer(
model=args.model, # a model name, so TRL loads it after the DeepSpeed config exists
args=config,
train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
result = trainer.train()
trainer.save_model("qwen3-8b-sft/final") # ZeRO-3 gathers the 16-bit weights for this
if trainer.is_world_process_zero():
tokens = [e["num_tokens"] for e in trainer.state.log_history if "num_tokens" in e][-1]
print(f"tokens_per_s = {tokens / result.metrics['train_runtime']:.0f}")
print(f"peak_reserved_gib = {torch.cuda.max_memory_reserved() / 2**30:.1f}")
The batch on eight GPUs is 1 per GPU x 4 accumulation steps x 8 GPUs = 32 conversations per optimizer step, which DeepSpeed's config docs define as train_batch_size. One epoch of Capybara's 15,806 conversations is 494 steps (our calculation: 15,806 / 32).
Launch it on one VM#
The DeepSpeed launcher and torchrun both start one process per GPU, and the Transformers docs list both. They also list Accelerate's launcher, but it ignores the deepspeed setting in the script's config and needs a config file of its own, so this guide uses the first two:
# ZeRO-3 on every GPU in the VM
deepspeed --num_gpus 8 sft_ds.py --ds-config ds_zero3.json
# the same job with torchrun
torchrun --standalone --nproc-per-node=gpu sft_ds.py --ds-config ds_zero3.json
# ZeRO-2 on only GPUs 0 to 3
deepspeed --include localhost:0,1,2,3 sft_ds.py --ds-config ds_zero2.json
Start with --max-steps 50 to check memory and speed before you commit to the full epoch. A run with --max-steps writes no checkpoints, because the Trainer's timer would otherwise include saving about 115 GB of DeepSpeed state at the end (our calculation: 16.4 GB of bf16 weights plus 98.3 GB of fp32 optimizer state). The script prints tokens_per_s and rank 0's peak memory at the end, the Trainer logs the loss every 10 steps, and the cost guide turns tokens per second into hours and dollars. The DDP and FSDP guide shows how to check the links between GPUs with nvidia-smi topo -m and an all-reduce test, which matters most for ZeRO-3 on PCIe-only VMs such as the L40S.
Checkpoints and getting the weights out#
Every save_steps, and once more when training ends, the Trainer writes a checkpoint-N folder with DeepSpeed's sharded state, one set of files per rank. The optimizer part alone is 12 bytes per parameter, about 98 GB for Qwen3-8B (our calculation), which is why save_total_limit=2 keeps only the newest two. DeepSpeed also drops a zero_to_fp32.py script into each checkpoint, which rebuilds full fp32 weights without GPUs:
cd qwen3-8b-sft/checkpoint-400
python zero_to_fp32.py . ../fp32-weights --safe_serialization
For deployment you rarely need that, because stage3_gather_16bit_weights_on_model_save makes trainer.save_model write a normal 16-bit Hugging Face model to qwen3-8b-sft/final. DeepSpeed's docs warn that gathering the weights onto one GPU is slow and memory hungry, so do it once at the end.
Stopping a QuantaCloud instance deletes its disk, and nothing on it survives. Push the output to a private Hugging Face repo from a second terminal, every 30 minutes, while training runs:
hf auth login
hf upload your-username/qwen3-8b-sft ./qwen3-8b-sft . --private --every=30
The file transfer guide covers rsync and rclone to your own storage instead. Before a long run, fund it or turn on auto top-up: the first hour is charged at launch, each further hour when the previous one is used up, and if the balance cannot cover the next hour the instance is terminated with its disk (pricing).
When to use FSDP instead#
I would use FSDP2 for new code and keep DeepSpeed for the jobs where its offload earns its keep. The two overlap more than their names suggest: Accelerate's comparison maps FSDP's full sharding to ZeRO stage 3, and both run on the same VM with the same launchers. The differences that decide it:
| Question | DeepSpeed ZeRO | FSDP2 (fully_shard) |
|---|---|---|
| Extra software | A separate package whose ops compile on first use | Part of PyTorch |
| Offload | Optimizer and weights separately, to CPU or NVMe | All or nothing, to CPU |
| Master weights | fp32 by default, and Accelerate warns the upcast can cost memory on few GPUs | The dtype the model is loaded in, fp32 in the DDP and FSDP guide's script |
| Saving a 16-bit model | One config flag gathers it at save time | Gather a full state dict with PyTorch's checkpoint APIs |
| Tool support | Transformers, TRL, Accelerate, Axolotl | The same tools. Axolotl recommends FSDP2 for new users |
Use DeepSpeed when a model only fits with optimizer offload, when you want NVMe offload, or when the recipe you are following already ships a DeepSpeed config. Use FSDP2 when the model fits sharded across the VM's GPUs and you want one less dependency. For LoRA and QLoRA on one VM, plain DDP is often enough, and Axolotl switches between DDP, DeepSpeed and FSDP with one line of YAML.
FAQ#
What is DeepSpeed?
DeepSpeed is an open-source library for training large models on many GPUs. Its best-known part is ZeRO, the Zero Redundancy Optimizer, which removes the duplicate copies of training state that plain data parallelism keeps on every GPU. It plugs into the Hugging Face Trainer through a JSON config, as above.
Which ZeRO stage should I start with?
The lowest one that fits. Axolotl's docs say to go from stage 1 to 2 to 3, and the Transformers docs say to use ZeRO-3 only when the model does not fit with ZeRO-2. Stages 1 and 2 cost no extra communication, and stage 3 costs up to 1.5 times as much.
Does ZeRO-3 need NVLink?
No, but it moves the most data of any stage, so it gains the most from a fast link. On a PCIe-only VM such as 8x L40S, ZeRO-3 still runs, and the extra traffic is the price of fitting the model. Check the links with nvidia-smi topo -m, and read SXM vs PCIe for what each link between GPUs carries.
Can DeepSpeed train across several VMs?
The launcher supports multiple nodes with a hostfile, but on-demand QuantaCloud VMs are single machines of up to eight GPUs, and multi-node training is not self-serve. For jobs that need more, we order and build GPU clusters to your spec, with InfiniBand between nodes, and quote the configuration, lead time and terms in writing: send a capacity brief.
The rule I follow for ZeRO: run the estimator for your model and GPU count, take the lowest stage that leaves room for activations, and add offload only when nothing else fits. For Qwen3-8B that means ZeRO-2 on an 8x A100 SXM4 and ZeRO-3 on an 8x L40S. If the model fits sharded without offload, FSDP2 does the same job inside PyTorch, and the fine-tuning overview links every other way to train.
Launch 8x L40S