GPU guide

Check your GPU, NVIDIA driver and CUDA version

nvidia-smi shows the newest CUDA your driver supports, not what is installed. Check driver, toolkit and PyTorch versions, then fix the usual mismatches.

Faiz Ahmed11 min read

The honest answer is that a GPU machine has three CUDA versions, and each command reports a different one. nvidia-smi shows the newest CUDA the driver supports, nvcc --version shows the toolkit you compile with, and torch.version.cuda shows the CUDA your PyTorch build was made for. They are allowed to differ. What has to hold is that the driver is new enough for everything you run: builds made with CUDA 13 need driver 580 or newer, and builds made with CUDA 12 need 525 or newer.

On QuantaCloud the driver comes with the instance. Every template runs on the same Ubuntu 22.04 VM image with the NVIDIA driver installed, and the deploy page offers a single OS image, so the driver is not something you pick. The job is to read it once, then choose PyTorch wheels, containers and toolkits that fit it. Connect to the instance over SSH and run the commands below.

Launch an RTX A6000 to check it yourself

Run three commands first#

The fastest check is three commands on the instance:

Terminal
nvidia-smi
nvcc --version
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"
CommandWhat it reportsWhere the number comes from
nvidia-smiThe GPU, the driver version and the newest CUDA that driver supportsThe NVIDIA driver
nvcc --versionThe release of the CUDA compilerA CUDA toolkit, if one is installed
torch.version.cudaThe CUDA version your PyTorch build was compiled forThe PyTorch wheel, which ships its own CUDA libraries
torch.cuda.is_available()Whether that build can use this driver and GPUBoth of the above

The third command needs PyTorch, and the PyTorch section below shows which build to install. PyTorch does not need nvcc: its wheels bring their own CUDA runtime, so nvcc: command not found next to a working torch.cuda.is_available() is normal. For scripts and logs, ask nvidia-smi for exact fields:

Terminal
nvidia-smi --query-gpu=name,driver_version,compute_cap,memory.total --format=csv

What the CUDA Version in nvidia-smi means#

The CUDA Version in nvidia-smi is a ceiling, not an install. NVIDIA's nvidia-smi manual describes it as "the latest CUDA version supported by the driver", and NVIDIA's Container Toolkit docs show this header from one of their machines:

Output
| NVIDIA-SMI 535.86.10    Driver Version: 535.86.10    CUDA Version: 12.2     |

That driver runs programs built with CUDA up to 12.2, with no CUDA toolkit installed at all. NVIDIA's current manual lists the same two fields as KMD Version and CUDA UMD Version and marks the old names deprecated, so newer drivers may print the new labels. The meaning is the same.

Two rules from NVIDIA's compatibility guide decide what runs. Within a major version, a build made with a newer toolkit still runs on an older driver with a reduced feature set, as long as the driver is 525 or newer for CUDA 12.x and 580 or newer for CUDA 13.x. Code that compiles PTX at run time is the exception and needs the newer driver. Across major versions nothing runs: a CUDA 13 build fails on a 570 driver. NVIDIA's forward-compatibility package lifts that limit only on data-center GPUs, select NGC Server Ready RTX cards and Jetson boards.

The rest of the output reads the same on every GPU:

FieldWhat it means
Persistence-MWhether persistence mode is on
PerfThe performance state, from P0 (maximum performance) to P12 (minimum)
Pwr:Usage/CapThe last measured power draw of the board in watts, next to its power limit
Memory-UsageUsed and total GPU memory in MiB
GPU-UtilThe share of the last sample period in which at least one kernel was running, not how much of the GPU was busy
Volatile Uncorr. ECCUncorrectable memory errors counted since the driver loaded
Compute M.Whether one or several compute applications may use the GPU at once
MIG M.Whether the GPU is split into MIG instances
ProcessesThe processes using the GPU and the memory each one holds

What QuantaCloud instances report#

The one thing I always check is the driver on the instance I am paying for. The Bare Metal image is not version-pinned, so run nvidia-smi on your own instance.

On Linux, Blackwell GPUs work only with NVIDIA's open kernel modules, so an RTX PRO 6000 instance that lists its GPU in nvidia-smi is already running them.

Minimum drivers for each CUDA release and GPU#

Each CUDA release has a minimum driver, and each GPU needs a minimum CUDA release. These are NVIDIA's Linux minimums:

CUDA toolkitReleasedMinimum Linux driverNotes
11.8October 2022520.61.05First release with Hopper and Ada
12.8January 2025570.26First release with Blackwell (sm_100, sm_120)
12.9May 2025575.51.03Adds sm_103 for B300
13.0August 2025580.65.06Drops Maxwell, Pascal and Volta
13.1December 2025590.44.01
13.2March 2026595.45.04
13.3May 2026610.43.02
13.4September 2026R615 branchThe Linux driver no longer ships with the toolkit

For the GPUs QuantaCloud offers, that works out as follows:

GPUCompute capabilityCUDA releases that support itDriver that CUDA release needs
RTX A60008.6Every current 12.x and 13.x release525 for CUDA 12, 580 for CUDA 13
A100 80GB (SXM4 and PCIe)8.0Every current 12.x and 13.x release525 for CUDA 12, 580 for CUDA 13
L40, L40S, RTX 6000 Ada8.911.8 and later520.61.05 for CUDA 11.8
H100 PCIe, H200 NVL9.011.8 and later520.61.05 for CUDA 11.8
RTX PRO 6000 Blackwell12.012.8 and later570.26 for CUDA 12.8

The last column is the CUDA side of the rule. A GPU model that shipped after its architecture's first CUDA release, such as the H200 NVL, also needs a driver branch that knows the card. On a rented VM that part is already settled: if nvidia-smi lists the GPU, the driver supports it.

Pick PyTorch to match the driver#

The rule I follow on a rented VM is to choose PyTorch from the driver, never the driver from PyTorch. Since PyTorch 2.11, a plain pip install torch on Linux installs the CUDA 13.0 build, which needs driver 580 or newer. On an older driver, install a CUDA 12 build from PyTorch's own index:

CUDA Version in nvidia-smiInstallRTX PRO 6000 Blackwell
13.0 or newer (driver 580 or newer)uv pip install torchWorks
12.8 or 12.9 (driver 570 to 579)uv pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu129Works
Below 12.8 (driver 525 to 569)uv pip install torch --index-url https://download.pytorch.org/whl/cu126No: cu126 has no Blackwell kernels, and Blackwell needs driver 570 or newer anyway

On download.pytorch.org, Linux cu129 wheels stop at PyTorch 2.13.0 and cu128 wheels at 2.11.0, so with a driver from 570 to 579, PyTorch 2.13.0 is the newest release that runs on Blackwell. The commands assume a uv environment, which needs nothing from apt:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/venv --python 3.12
source ~/venv/bin/activate
uv pip install torch
python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"

Plain pip takes the same arguments. uv pip install torch --torch-backend=auto picks the index for you from the installed driver, and on Blackwell it is still worth confirming the architecture list afterwards:

Terminal
python -c "import torch; print(torch.cuda.get_device_capability(0), torch.cuda.get_arch_list())"

On an RTX PRO 6000 the capability is (12, 0) and the list must include sm_120.

The same split runs through the rest of the stack. vLLM 0.30.0's default wheel and latest image are CUDA 13.0 builds that need driver 580 or newer, and v0.30.0-cu129 is the image for older drivers, as deploying vLLM with Docker and installing vLLM show. ComfyUI requires a cu130 PyTorch build on RTX 20-series and newer GPUs, which again means driver 580 or newer (run ComfyUI on a cloud GPU). The PyTorch + Jupyter template ships its own PyTorch build without a pinned version, so run the torch.version.cuda check in a notebook cell before you install anything on top.

Install the CUDA toolkit only when you compile#

You need nvcc only when something compiles CUDA code on the instance: flash-attn built from source, a PyTorch C++ or CUDA extension, or your own kernels. Pick a toolkit no newer than the driver's CUDA Version and install it from NVIDIA's repository:

  1. Add NVIDIA's CUDA repository for Ubuntu 22.04:

    Terminal
    wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
    sudo dpkg -i cuda-keyring_1.1-1_all.deb
    sudo apt-get update
    
  2. Install the versioned toolkit, here 12.8 for a driver that reports 12.8:

    Terminal
    sudo apt-get install -y cuda-toolkit-12-8
    
  3. Put it on your PATH and check it:

    Terminal
    export PATH=/usr/local/cuda-12.8/bin:$PATH
    nvcc --version
    

A cuda-toolkit-X-Y package never upgrades past its X.Y series, and NVIDIA's Ubuntu 22.04 repository carries versioned toolkits from 11.7 to 13.4. Leave cuda-drivers alone: it installs a driver, and the instance already has one. If the shell suggests sudo apt install nvidia-cuda-toolkit, skip that too, because Ubuntu 22.04's own package is CUDA 11.5.1, older than the first release with Hopper support. The toolkit sits on the instance's disk, so it is gone after a stop: keep these lines in your setup script.

Common mismatch errors and their fixes#

Each of these errors says which side is too old or too new:

ErrorWhat it meansFix
The NVIDIA driver on your system is too old (found version 12080)PyTorch was built for a newer CUDA than the driver supports. The number is the driver's CUDA version: 12080 is 12.8Install the PyTorch build for your driver from the table above
CUDA driver version is insufficient for CUDA runtime versionA program was built for a newer CUDA than the driver allows, such as a CUDA 13 build on a 570 driverInstall or build a version for the driver's CUDA
with CUDA capability sm_120 is not compatible with the current PyTorch installationA PyTorch build without Blackwell kernels, such as cu126, on an RTX PRO 6000Install a cu128, cu129, cu130 or cu132 build and check for sm_120 in torch.cuda.get_arch_list()
no kernel image is available for execution on the deviceThe binary has no code for this GPU's architectureUse a build that targets your compute capability
Torch not compiled with CUDA enabledA CPU-only PyTorch wheelReinstall PyTorch from a CUDA index
nvcc: command not foundNo CUDA toolkit is installed, and the driver does not include oneInstall cuda-toolkit-X-Y only if you compile
nvidia-smi: command not found inside a containerThe container was started without GPU accessStart it with --gpus all
could not select device driver "" with capabilities: [[gpu]]Docker has no NVIDIA runtime configuredInstall the NVIDIA Container Toolkit, run sudo nvidia-ctk runtime configure --runtime=docker and restart Docker
unsatisfied condition: cuda>=13.0, please update your driver to a newer version, or use an earlier cuda containerThe image needs a newer driver than the instance hasUse an image built for your driver's CUDA, such as vLLM's -cu129 tags

Check the GPU from inside Docker#

NVIDIA's sample command is the quickest container check:

Terminal
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

It prints the instance's driver and CUDA Version even though the ubuntu image contains no NVIDIA software, because the NVIDIA runtime mounts the driver into the container. The CUDA inside any other image is whatever it was built with, and it still has to fit that driver.

When the command fails, running Docker with a GPU and deploying vLLM with Docker cover the NVIDIA Container Toolkit setup.


My order on a fresh instance never changes: nvidia-smi first, then PyTorch, containers and toolkits chosen to fit its CUDA number. On Blackwell, check that every part of your stack has a CUDA 12.8 or newer build before you launch an RTX PRO 6000. If the versions line up and the job still fails with out of memory, that is a sizing problem rather than a driver problem, and fixing CUDA out of memory is the next step.

Launch an RTX PRO 6000 Blackwell

Keep building

Choose your next step.