The honest answer is that a GPU machine has three CUDA versions, and each command reports a different one. nvidia-smi shows the newest CUDA the driver supports, nvcc --version shows the toolkit you compile with, and torch.version.cuda shows the CUDA your PyTorch build was made for. They are allowed to differ. What has to hold is that the driver is new enough for everything you run: builds made with CUDA 13 need driver 580 or newer, and builds made with CUDA 12 need 525 or newer.
On QuantaCloud the driver comes with the instance. Every template runs on the same Ubuntu 22.04 VM image with the NVIDIA driver installed, and the deploy page offers a single OS image, so the driver is not something you pick. The job is to read it once, then choose PyTorch wheels, containers and toolkits that fit it. Connect to the instance over SSH and run the commands below.
Launch an RTX A6000 to check it yourselfRun three commands first#
The fastest check is three commands on the instance:
nvidia-smi
nvcc --version
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"
| Command | What it reports | Where the number comes from |
|---|---|---|
nvidia-smi | The GPU, the driver version and the newest CUDA that driver supports | The NVIDIA driver |
nvcc --version | The release of the CUDA compiler | A CUDA toolkit, if one is installed |
torch.version.cuda | The CUDA version your PyTorch build was compiled for | The PyTorch wheel, which ships its own CUDA libraries |
torch.cuda.is_available() | Whether that build can use this driver and GPU | Both of the above |
The third command needs PyTorch, and the PyTorch section below shows which build to install. PyTorch does not need nvcc: its wheels bring their own CUDA runtime, so nvcc: command not found next to a working torch.cuda.is_available() is normal. For scripts and logs, ask nvidia-smi for exact fields:
nvidia-smi --query-gpu=name,driver_version,compute_cap,memory.total --format=csv
What the CUDA Version in nvidia-smi means#
The CUDA Version in nvidia-smi is a ceiling, not an install. NVIDIA's nvidia-smi manual describes it as "the latest CUDA version supported by the driver", and NVIDIA's Container Toolkit docs show this header from one of their machines:
| NVIDIA-SMI 535.86.10 Driver Version: 535.86.10 CUDA Version: 12.2 |
That driver runs programs built with CUDA up to 12.2, with no CUDA toolkit installed at all. NVIDIA's current manual lists the same two fields as KMD Version and CUDA UMD Version and marks the old names deprecated, so newer drivers may print the new labels. The meaning is the same.
Two rules from NVIDIA's compatibility guide decide what runs. Within a major version, a build made with a newer toolkit still runs on an older driver with a reduced feature set, as long as the driver is 525 or newer for CUDA 12.x and 580 or newer for CUDA 13.x. Code that compiles PTX at run time is the exception and needs the newer driver. Across major versions nothing runs: a CUDA 13 build fails on a 570 driver. NVIDIA's forward-compatibility package lifts that limit only on data-center GPUs, select NGC Server Ready RTX cards and Jetson boards.
The rest of the output reads the same on every GPU:
| Field | What it means |
|---|---|
| Persistence-M | Whether persistence mode is on |
| Perf | The performance state, from P0 (maximum performance) to P12 (minimum) |
| Pwr:Usage/Cap | The last measured power draw of the board in watts, next to its power limit |
| Memory-Usage | Used and total GPU memory in MiB |
| GPU-Util | The share of the last sample period in which at least one kernel was running, not how much of the GPU was busy |
| Volatile Uncorr. ECC | Uncorrectable memory errors counted since the driver loaded |
| Compute M. | Whether one or several compute applications may use the GPU at once |
| MIG M. | Whether the GPU is split into MIG instances |
| Processes | The processes using the GPU and the memory each one holds |
What QuantaCloud instances report#
The one thing I always check is the driver on the instance I am paying for. The Bare Metal image is not version-pinned, so run nvidia-smi on your own instance.
On Linux, Blackwell GPUs work only with NVIDIA's open kernel modules, so an RTX PRO 6000 instance that lists its GPU in nvidia-smi is already running them.
Minimum drivers for each CUDA release and GPU#
Each CUDA release has a minimum driver, and each GPU needs a minimum CUDA release. These are NVIDIA's Linux minimums:
| CUDA toolkit | Released | Minimum Linux driver | Notes |
|---|---|---|---|
| 11.8 | October 2022 | 520.61.05 | First release with Hopper and Ada |
| 12.8 | January 2025 | 570.26 | First release with Blackwell (sm_100, sm_120) |
| 12.9 | May 2025 | 575.51.03 | Adds sm_103 for B300 |
| 13.0 | August 2025 | 580.65.06 | Drops Maxwell, Pascal and Volta |
| 13.1 | December 2025 | 590.44.01 | |
| 13.2 | March 2026 | 595.45.04 | |
| 13.3 | May 2026 | 610.43.02 | |
| 13.4 | September 2026 | R615 branch | The Linux driver no longer ships with the toolkit |
For the GPUs QuantaCloud offers, that works out as follows:
| GPU | Compute capability | CUDA releases that support it | Driver that CUDA release needs |
|---|---|---|---|
| RTX A6000 | 8.6 | Every current 12.x and 13.x release | 525 for CUDA 12, 580 for CUDA 13 |
| A100 80GB (SXM4 and PCIe) | 8.0 | Every current 12.x and 13.x release | 525 for CUDA 12, 580 for CUDA 13 |
| L40, L40S, RTX 6000 Ada | 8.9 | 11.8 and later | 520.61.05 for CUDA 11.8 |
| H100 PCIe, H200 NVL | 9.0 | 11.8 and later | 520.61.05 for CUDA 11.8 |
| RTX PRO 6000 Blackwell | 12.0 | 12.8 and later | 570.26 for CUDA 12.8 |
The last column is the CUDA side of the rule. A GPU model that shipped after its architecture's first CUDA release,
such as the H200 NVL, also needs a driver branch that knows the card. On a rented VM that part is already settled: if
nvidia-smi lists the GPU, the driver supports it.
Pick PyTorch to match the driver#
The rule I follow on a rented VM is to choose PyTorch from the driver, never the driver from PyTorch. Since PyTorch 2.11, a plain pip install torch on Linux installs the CUDA 13.0 build, which needs driver 580 or newer. On an older driver, install a CUDA 12 build from PyTorch's own index:
CUDA Version in nvidia-smi | Install | RTX PRO 6000 Blackwell |
|---|---|---|
| 13.0 or newer (driver 580 or newer) | uv pip install torch | Works |
| 12.8 or 12.9 (driver 570 to 579) | uv pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu129 | Works |
| Below 12.8 (driver 525 to 569) | uv pip install torch --index-url https://download.pytorch.org/whl/cu126 | No: cu126 has no Blackwell kernels, and Blackwell needs driver 570 or newer anyway |
On download.pytorch.org, Linux cu129 wheels stop at PyTorch 2.13.0 and cu128 wheels at 2.11.0, so with a driver from 570 to 579, PyTorch 2.13.0 is the newest release that runs on Blackwell. The commands assume a uv environment, which needs nothing from apt:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv venv ~/venv --python 3.12
source ~/venv/bin/activate
uv pip install torch
python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"
Plain pip takes the same arguments. uv pip install torch --torch-backend=auto picks the index for you from the installed driver, and on Blackwell it is still worth confirming the architecture list afterwards:
python -c "import torch; print(torch.cuda.get_device_capability(0), torch.cuda.get_arch_list())"
On an RTX PRO 6000 the capability is (12, 0) and the list must include sm_120.
The same split runs through the rest of the stack. vLLM 0.30.0's default wheel and latest image are CUDA 13.0 builds that need driver 580 or newer, and v0.30.0-cu129 is the image for older drivers, as deploying vLLM with Docker and installing vLLM show. ComfyUI requires a cu130 PyTorch build on RTX 20-series and newer GPUs, which again means driver 580 or newer (run ComfyUI on a cloud GPU). The PyTorch + Jupyter template ships its own PyTorch build without a pinned version, so run the torch.version.cuda check in a notebook cell before you install anything on top.
Install the CUDA toolkit only when you compile#
You need nvcc only when something compiles CUDA code on the instance: flash-attn built from source, a PyTorch C++ or CUDA extension, or your own kernels. Pick a toolkit no newer than the driver's CUDA Version and install it from NVIDIA's repository:
-
Add NVIDIA's CUDA repository for Ubuntu 22.04:
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb sudo dpkg -i cuda-keyring_1.1-1_all.deb sudo apt-get update -
Install the versioned toolkit, here 12.8 for a driver that reports 12.8:
sudo apt-get install -y cuda-toolkit-12-8 -
Put it on your PATH and check it:
export PATH=/usr/local/cuda-12.8/bin:$PATH nvcc --version
A cuda-toolkit-X-Y package never upgrades past its X.Y series, and NVIDIA's Ubuntu 22.04 repository carries versioned toolkits from 11.7 to 13.4. Leave cuda-drivers alone: it installs a driver, and the instance already has one. If the shell suggests sudo apt install nvidia-cuda-toolkit, skip that too, because Ubuntu 22.04's own package is CUDA 11.5.1, older than the first release with Hopper support. The toolkit sits on the instance's disk, so it is gone after a stop: keep these lines in your setup script.
Common mismatch errors and their fixes#
Each of these errors says which side is too old or too new:
| Error | What it means | Fix |
|---|---|---|
The NVIDIA driver on your system is too old (found version 12080) | PyTorch was built for a newer CUDA than the driver supports. The number is the driver's CUDA version: 12080 is 12.8 | Install the PyTorch build for your driver from the table above |
CUDA driver version is insufficient for CUDA runtime version | A program was built for a newer CUDA than the driver allows, such as a CUDA 13 build on a 570 driver | Install or build a version for the driver's CUDA |
with CUDA capability sm_120 is not compatible with the current PyTorch installation | A PyTorch build without Blackwell kernels, such as cu126, on an RTX PRO 6000 | Install a cu128, cu129, cu130 or cu132 build and check for sm_120 in torch.cuda.get_arch_list() |
no kernel image is available for execution on the device | The binary has no code for this GPU's architecture | Use a build that targets your compute capability |
Torch not compiled with CUDA enabled | A CPU-only PyTorch wheel | Reinstall PyTorch from a CUDA index |
nvcc: command not found | No CUDA toolkit is installed, and the driver does not include one | Install cuda-toolkit-X-Y only if you compile |
nvidia-smi: command not found inside a container | The container was started without GPU access | Start it with --gpus all |
could not select device driver "" with capabilities: [[gpu]] | Docker has no NVIDIA runtime configured | Install the NVIDIA Container Toolkit, run sudo nvidia-ctk runtime configure --runtime=docker and restart Docker |
unsatisfied condition: cuda>=13.0, please update your driver to a newer version, or use an earlier cuda container | The image needs a newer driver than the instance has | Use an image built for your driver's CUDA, such as vLLM's -cu129 tags |
Check the GPU from inside Docker#
NVIDIA's sample command is the quickest container check:
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
It prints the instance's driver and CUDA Version even though the ubuntu image contains no NVIDIA software, because the NVIDIA runtime mounts the driver into the container. The CUDA inside any other image is whatever it was built with, and it still has to fit that driver.
When the command fails, running Docker with a GPU and deploying vLLM with Docker cover the NVIDIA Container Toolkit setup.
My order on a fresh instance never changes: nvidia-smi first, then PyTorch, containers and toolkits chosen to fit its CUDA number. On Blackwell, check that every part of your stack has a CUDA 12.8 or newer build before you launch an RTX PRO 6000. If the versions line up and the job still fails with out of memory, that is a sizing problem rather than a driver problem, and fixing CUDA out of memory is the next step.