GPU guide

Why GPU utilization is low during training, and how to fix it

Why GPU utilization is low in training: measure it with nvidia-smi, dmon and the PyTorch profiler, then fix data loading, small batches and sync points.

Faiz Ahmed11 min read

The usual reason GPU utilization is low is that the GPU is waiting: for the next batch from the data loader, for the CPU to launch its next small kernel, or for the other GPUs to finish exchanging gradients. The number you see in nvidia-smi, GPU-Util, is the share of time at least one kernel was running, so a low value is idle time on a GPU you pay for by the hour, while a high value does not prove the GPU is doing much. Measure first with nvidia-smi and the PyTorch profiler, then fix the input pipeline, because that is the cause I would check first.

What GPU-Util measures, and what it misses#

NVIDIA's manual defines GPU utilization as the percent of time over the past sample period during which one or more kernels was executing on the GPU, with a sample period between 1 second and 1/6 of a second depending on the product. Memory utilization is the same idea for memory: the share of time global memory was being read or written. Neither says how much of the GPU a kernel used. One small kernel that keeps a single SM busy all the time reads as 100%.

The deeper counters answer that question. NVIDIA's DCGM defines SM Activity as the fraction of time at least one warp was active, averaged over all SMs, and says a value of 0.8 or more is necessary but not sufficient for effective use, while below 0.5 likely means ineffective use. On Hopper and newer GPUs, which on QuantaCloud means the H100 PCIe, H200 NVL and RTX PRO 6000, NVML's GPM counters report similar figures without installing DCGM, and nvidia-smi dmon --gpm-metrics 2,5,10 prints SM utilization, tensor-core utilization and memory bandwidth utilization.

The cost of a low number is simple arithmetic. At 40% GPU-Util, no kernel is running 60% of the time. On an H200 NVL at $3.43 an hour on 2026-09-27 (the console price now), a 10-hour run at 40% pays $20.58 for time the GPU spent waiting, and lifting it to 90% would finish the same work in about 4.4 hours, if the kernels themselves take as long as before (our calculation).

Measure it in three steps#

Measure during a steady stretch of training, not during the first minute, when the model loads and the first steps compile.

  1. Log utilization, memory and power once a second from a second SSH session:

    Terminal
    nvidia-smi --query-gpu=utilization.gpu,utilization.memory,memory.used,power.draw --format=csv -l 1
    
  2. Watch the pattern with dmon, which prints one line per second per GPU:

    Terminal
    nvidia-smi dmon -s pu -d 1
    

    The sm and mem columns are the utilization figures and pwr is the power draw. A GPU that is fed well stays high and steady. A starved one alternates between bursts and near zero.

  3. Find where the time goes with the PyTorch profiler, described in the next section.

NVIDIA's manual describes dmon for bare-metal Linux, and dmon prints a dash for any metric it cannot read. If it shows dashes, the --query-gpu loop in step 1 still gives you GPU-Util.

Check the CPU side at the same time. top or htop showing one Python process pinned at 100% of a single core, with the other cores idle, is the signature of data loading in the main process.

Profile the training loop#

The PyTorch profiler records both CPU and GPU activity, and its schedule keeps the trace small on a long run. Wrap a few steps of the loop:

Python
import torch
from torch.profiler import profile, schedule, ProfilerActivity

def report(prof):
    print(prof.key_averages().table(sort_by="self_cuda_time_total", row_limit=15))
    prof.export_chrome_trace(f"trace_{prof.step_num}.json")

with profile(
    activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
    schedule=schedule(wait=5, warmup=2, active=5, repeat=1),
    on_trace_ready=report,
) as prof:
    for step, (x, y) in enumerate(loader):
        train_step(x, y)  # your forward, backward and optimizer step
        prof.step()
        if step >= 12:
            break

The table lists the operations that took the most GPU time. Sort by self_cpu_time_total as well: a large row named enumerate(DataLoader)#_SingleProcessDataLoaderIter.__next__ is the training loop waiting for its next batch. Copy the trace file to your laptop and open it in Chrome's trace viewer at chrome://tracing. Gaps in the GPU stream between kernels are the idle time, and the CPU rows above each gap show what the process was doing instead.

Causes and fixes#

Those three readings point to one of a short list of causes:

What you seeLikely causeFix
GPU-Util drops towards zero between steps, one CPU core at 100%, the DataLoader row is largeData loading in the main process (num_workers=0, the default)num_workers above 0, pin_memory=True, persistent_workers=True
Every vCPU busy and the GPU still waitsDecoding and augmentation cost more CPU than the VM hasPreprocess once and store tensors, move augmentation to the GPU, or pick a VM with more vCPUs per GPU
Utilization high but SM activity low, many tiny kernels in the traceSmall batches or a small model, so kernel launches dominateLarger batches, mixed precision, torch.compile(model, mode="reduce-overhead")
Regular dips, with synchronize calls in the traceSync points in the loop, such as .item() or printing a CUDA tensor every stepLog every N steps and keep metrics on the GPU until then
Dips at every epoch boundaryDataLoader workers restart each epochpersistent_workers=True
Several GPUs, with all-reduce kernels taking a large shareGradient exchange over a slow linkCheck nvidia-smi topo -m, accumulate gradients with no_sync(), use larger batches per GPU
Long stretches at zeroNothing is running: an idle notebook, evaluation on the CPU, a checkpoint being writtenStop the instance when the work is done

Fix the input pipeline first#

The input pipeline comes first because PyTorch's DataLoader loads data in the main process by default, so the GPU waits while every batch is read and transformed. Worker processes let loading overlap with training, and pinned memory lets the copy to the GPU run asynchronously:

Python
loader = DataLoader(
    dataset,
    batch_size=256,
    shuffle=True,
    num_workers=11,           # up to the vCPUs you have, minus one for the main process
    pin_memory=True,
    persistent_workers=True,  # keep workers alive between epochs
    prefetch_factor=4,        # batches queued per worker, 2 by default
)

for x, y in loader:
    x = x.cuda(non_blocking=True)
    y = y.cuda(non_blocking=True)

The vCPU count caps num_workers, and it differs by GPU. These are the single-GPU configurations in the 2026-09-27 catalog, and the offer card shows the current figure before you launch:

1x GPUvCPUsRAM
RTX A60006 to 1224 to 64 GB
RTX 6000 Ada, L40S1272 GB
L401472 GB
A100 SXM414 to 16100 to 120 GB
RTX PRO 6000, H200 NVL16144 and 180 GB
H100 PCIe20128 GB

PyTorch warns when you ask for more workers than it finds CPUs: This DataLoader will create N worker processes in total. Our suggested max number of worker in current system is M. Heed it, because extra workers only compete for the same cores. If the loader still cannot keep up with every core busy, the fix is less CPU work per sample: decode and resize a dataset once and store the tensors, or move augmentation onto the GPU. Keep the data on the VM's local disk too, not on a network share you read during training. The file transfer guide covers getting it there.

Give the GPU bigger pieces of work#

Once the data arrives in time, small kernels are the next suspect. A small batch on a large GPU finishes each kernel so quickly that the CPU cannot launch the next one fast enough. Raise the batch size as far as memory allows, and use mixed precision with torch.autocast("cuda", dtype=torch.bfloat16), which runs on the Tensor Cores and, according to PyTorch's tuning guide, gives up to 3 times the speed on Volta and newer GPUs. For small batches that cannot grow, torch.compile(model, mode="reduce-overhead") cuts the launch overhead with CUDA graphs, at some extra memory, though PyTorch notes that it does not work for every model. For convolutional networks with fixed input sizes, set torch.backends.cudnn.benchmark = True. If the larger batch runs out of memory, the CUDA out of memory guide covers the trade-offs.

Remove sync points from the loop#

A sync point makes the CPU wait until the GPU has finished everything queued, so the queue runs dry. PyTorch's tuning guide lists the usual ones: print(cuda_tensor), cuda_tensor.item(), .cpu() and .to() copies, cuda_tensor.nonzero(), and Python if statements on CUDA values. A common one in training code is loss.item() on every step for logging. Accumulate the loss on the GPU and call .item() every 50 or 100 steps instead.

Multi-GPU jobs wait on each other#

With several GPUs, every data-parallel step ends with an all-reduce of the gradients, and slow links make every GPU wait. Check the link with nvidia-smi topo -m, as the SXM, PCIe and NVLink guide explains. NCCL copies directly between GPUs over NVLink or PCIe when peer-to-peer access works, and goes through shared host memory when it does not, and NCCL_DEBUG=INFO turns on its own log for a closer look. With gradient accumulation, DDP's no_sync() skips the all-reduce on every step but the last of each accumulation cycle, and a larger batch per GPU spreads the exchange over more work. The DDP and FSDP guide sets up the job itself.

A worked example on an RTX A6000#

This script trains ResNet-50 on images it generates on the CPU, so it needs no dataset download, and each 500 x 375 image goes through the transforms of a standard ImageNet training pipeline: a random resized crop, a flip and normalization. It prints the images per second for the number of DataLoader workers you pass. Launch the Bare Metal template, install PyTorch with uv, which picks the build that matches the driver, and save the script as util_demo.py:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv --python 3.12 --seed --managed-python ~/util
source ~/util/bin/activate
uv pip install torch torchvision --torch-backend=auto
Python
import sys, time, torch, torchvision
from torch.utils.data import DataLoader
from torchvision import transforms
from torchvision.datasets import FakeData

workers = int(sys.argv[1])
tf = transforms.Compose([
    transforms.RandomResizedCrop(224),
    transforms.RandomHorizontalFlip(),
    transforms.ToTensor(),
    transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
data = FakeData(size=128 * 130, image_size=(3, 375, 500), transform=tf)
loader = DataLoader(data, batch_size=128, shuffle=True, num_workers=workers,
                    pin_memory=True, persistent_workers=workers > 0)
model = torchvision.models.resnet50().cuda()
opt = torch.optim.SGD(model.parameters(), lr=0.1, momentum=0.9)
loss_fn = torch.nn.CrossEntropyLoss()

for step, (x, y) in enumerate(loader):
    if step == 20:
        torch.cuda.synchronize()
        start = time.time()
    x, y = x.cuda(non_blocking=True), y.cuda(non_blocking=True)
    with torch.autocast("cuda", dtype=torch.bfloat16):
        loss = loss_fn(model(x), y)
    opt.zero_grad(set_to_none=True)
    loss.backward()
    opt.step()
torch.cuda.synchronize()
print(f"{workers} workers: {(step - 19) * 128 / (time.time() - start):.0f} images per second")

Run it twice, with no workers and then with one worker per vCPU but one, while the nvidia-smi loop from the first step logs GPU-Util in a second session:

Terminal
python util_demo.py 0
python util_demo.py $(( $(nproc) - 1 ))

Stop paying for idle time#

The simplest idle time to remove is a GPU with nothing running. QuantaCloud bills from launch until you stop, and refunds the unused seconds of the current hour when you do, so an instance left open overnight after the job finished is pure cost. Stopping terminates the instance and deletes its disk, so copy checkpoints and logs off first. The auto-stop guide sets up a watchdog that stops an idle instance through the API, and the pricing page has the billing rules.

GPU utilization FAQ#

What is a good GPU utilization for training?

Close to 100% GPU-Util for the whole run, with SM activity of 0.8 or more where you can read it, which NVIDIA calls necessary, but not sufficient, for effective use. A training job that sits well below that is waiting on something, and the profiler shows what.

Why does nvidia-smi show 100% but training is still slow?

GPU-Util only counts the time at least one kernel was running. Small kernels that use a fraction of the GPU still read as 100%, so check SM activity, or the profiler's kernel list, before concluding that the GPU is the limit.

Why is my GPU utilization 0%?

Nothing is running on the GPU. Check that your code actually uses it: torch.cuda.is_available() should be True, and the model and tensors should be on cuda. Checking your GPU, driver and CUDA version covers a PyTorch build that cannot see the GPU.

Should I rent a bigger GPU when utilization is low?

Not while the input pipeline is the limit. A bigger GPU finishes each kernel sooner and then waits longer for the next batch, so utilization drops further. Fix the pipeline first, and move up a tier only when the GPU is busy and the job is still too slow.


My rule: measure before you change anything, fix the data loader before the model, and read the profiler before you rent a bigger GPU. A GPU kept close to 100% busy finishes a training run for the least money, because every hour you pay for is an hour of work. When the pipeline keeps up, pick the GPU by memory and price from the GPU catalog.

Launch a GPU

Keep building

Choose your next step.