GPU guide

Run Docker containers with NVIDIA GPUs

Install the NVIDIA Container Toolkit, pick GPUs with --gpus, choose CUDA images your driver runs, use GPUs in Compose, keep ports on 127.0.0.1, fix errors.

Faiz Ahmed10 min read

Docker needs three things to run a container on an NVIDIA GPU: the NVIDIA driver on the host, the NVIDIA Container Toolkit, and the --gpus flag on docker run. The container does not bring its own driver. The toolkit mounts the host driver into it when it starts, so the CUDA version an image was built with has to be one the host driver supports. On QuantaCloud, the Bare Metal template is an Ubuntu 22.04 VM with the NVIDIA driver and Docker already installed. What is left is to confirm the toolkit, pick images the driver can run, and keep the container's ports off the public internet.

Launch an RTX A6000 with Ubuntu and Docker

Check whether Docker can see the GPU#

The test NVIDIA publishes is a single container that runs nvidia-smi. Connect as ubuntu (the SSH guide has the setup) and run three checks in order:

  1. Read the driver on the host:

    Terminal
    nvidia-smi
    
  2. Check that ubuntu can talk to Docker:

    Terminal
    docker ps
    

    If it fails with permission denied while trying to connect to the docker API (Docker releases before 29 say the Docker daemon socket), add yourself to the docker group and log out and back in. Docker's docs warn that the group grants root-level privileges, which is fine on a single-user VM.

    Terminal
    sudo usermod -aG docker $USER
    
  3. Run the GPU inside a container:

    Terminal
    docker run --rm --gpus all ubuntu nvidia-smi
    

If the third command prints the same GPU table as the first, Docker and the GPU are working together, even though the ubuntu image contains no NVIDIA software. If Docker answers could not select device driver "" with capabilities: [[gpu]], or failed to discover GPU vendor from CDI: no known GPU vendor found from Docker 29.3 on, the toolkit is missing or not configured, and the next section fixes it.

Install the NVIDIA Container Toolkit#

The NVIDIA Container Toolkit is what lets Docker hand a GPU to a container. These are the steps from NVIDIA's install guide for Ubuntu, with the version it pinned on 2026-09-28:

  1. Install the prerequisites:

    Terminal
    sudo apt-get update && sudo apt-get install -y --no-install-recommends ca-certificates curl gnupg2
    
  2. Add NVIDIA's repository and its signing key:

    Terminal
    curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
    curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
      sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
      sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
    
  3. Install the toolkit packages:

    Terminal
    sudo apt-get update
    export NVIDIA_CONTAINER_TOOLKIT_VERSION=1.20.1-1
    sudo apt-get install -y \
      nvidia-container-toolkit=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \
      nvidia-container-toolkit-base=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \
      libnvidia-container-tools=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \
      libnvidia-container1=${NVIDIA_CONTAINER_TOOLKIT_VERSION}
    
  4. Register the NVIDIA runtime with Docker and restart it:

    Terminal
    sudo nvidia-ctk runtime configure --runtime=docker
    sudo systemctl restart docker
    
  5. Run the test from the previous section again:

    Terminal
    docker run --rm --gpus all ubuntu nvidia-smi
    

nvidia-ctk runtime configure writes the NVIDIA runtime into /etc/docker/daemon.json. The toolkit lives on the instance's disk, so it is gone after you stop the instance: the last section puts these steps in a script.

Choose GPUs with --gpus#

--gpus decides which GPUs a container sees. On a multi-GPU VM these are the forms you need:

FlagWhat the container gets
--gpus allEvery GPU on the VM
--gpus 2Two GPUs
--gpus device=0GPU 0 only, numbered as in nvidia-smi
--gpus '"device=0,2"'GPUs 0 and 2. The value needs both quotes because it contains a comma
--gpus device=GPU-UUIDOne GPU by the UUID nvidia-smi -L prints
--gpus 'all,"capabilities=compute,utility,video"'All GPUs, plus the video capability

By default a container gets the compute and utility parts of the driver, which cover CUDA and nvidia-smi. Video encoding and decoding with NVENC and NVDEC needs the video capability as well, as in the last row. Two other routes reach the same place. With --runtime=nvidia, the environment variable NVIDIA_VISIBLE_DEVICES=0,1 selects GPUs. And Docker Engine 28.2 and newer enable CDI by default, so docker run --rm --device nvidia.com/gpu=all ubuntu nvidia-smi works too: toolkit 1.18.0 and newer generate the device list themselves, and nvidia-ctk cdi list prints the names. From Docker 29.2, --gpus itself goes through CDI when the toolkit provides it.

For independent jobs on a 4-GPU VM, my rule is one container per GPU, each with its own --gpus device=N, so the jobs never compete for the same GPU's memory.

Pick a CUDA image the driver can run#

The rule I follow is to read the CUDA Version in nvidia-smi first and choose images built for a CUDA release the driver can run. CUDA 13.x builds need driver 580 or newer, and CUDA 12.x builds run on 525 or newer under NVIDIA's minor-version compatibility. NVIDIA's own images declare their CUDA version, and on the toolkit's older hook path a driver that is too old stops the container with unsatisfied condition: cuda>=13.0, please update your driver to a newer version, or use an earlier cuda container. The CDI path skips that check, so on Docker 29.2 and newer the container can start and the program inside fails instead, as PyTorch does with The NVIDIA driver on your system is too old.

CUDA Version in nvidia-smiCUDA base imagePyTorch imagevLLM image
13.0 or newer (driver 580 or newer)nvidia/cuda:13.0.2-base-ubuntu22.04pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtimevllm/vllm-openai:v0.30.0
12.8 or 12.9 (driver 570 to 579)nvidia/cuda:12.8.1-base-ubuntu22.04pytorch/pytorch:2.11.0-cuda12.8-cudnn9-runtimevllm/vllm-openai:v0.30.0-cu129

pytorch/pytorch:2.11.0-cuda12.8-cudnn9-runtime is the newest PyTorch image built on CUDA 12.8 or 12.9 on Docker Hub. On an RTX PRO 6000 Blackwell, stay on images built with CUDA 12.8 or newer: older builds have no Blackwell kernels and fail with no kernel image is available for execution on the device. NVIDIA's own PyTorch container, nvcr.io/nvidia/pytorch:26.08-py3, is built on CUDA 13.4.1, so it also needs driver 580 or newer. Checking the driver and CUDA version explains the numbers nvidia-smi prints.

NVIDIA's nvidia/cuda images come in three flavors: base has only the CUDA runtime, runtime adds the CUDA math libraries and NCCL, and devel adds the headers and compilers for building CUDA code. There is no latest tag, so docker pull nvidia/cuda fails with docker.io/nvidia/cuda:latest: not found, or manifest for nvidia/cuda:latest not found on Docker setups without the containerd image store. Name the full tag, and expect old tags to be removed over time under NVIDIA's support policy. Plan for the download too: pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtime is a 3.3 GB download and its devel variant 13.4 GB, and every new instance pulls them again.

A quick way to confirm a PyTorch image sees the GPU:

Terminal
docker run --rm --gpus all --ipc=host pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtime \
  python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"

GPUs in Docker Compose#

Compose asks for GPUs as a device reservation. This works in any current Compose version:

YAML
services:
  torch:
    image: pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtime
    command: python -c "import torch; print(torch.cuda.is_available(), torch.cuda.device_count())"
    ipc: host
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

Save it as compose.yaml and run docker compose up. The capabilities line is required, and Compose returns an error without it. Use count: 1 for one GPU, or device_ids: ['0', '3'] for specific ones, but not both in the same service. Docker Compose 2.30.0 and newer also accept a shorter form, gpus: all, in place of the whole deploy block. Check your version with docker compose version. If the command is missing, the plugin package is docker-compose-plugin when Docker came from Docker's own repository, and docker-compose-v2 with Ubuntu's docker.io package.

Open WebUI with Ollama and vLLM in Docker each include a full Compose file built on this, and ComfyUI in Docker runs ComfyUI the same way.

Publish ports on 127.0.0.1 only#

My rule for every GPU container is to publish its port on the loopback address and reach it through an SSH tunnel. Docker's own documentation calls publishing ports insecure by default: -p 8888:8888 listens on every interface of the VM, and Docker routes that traffic before ufw's rules see it, so the firewall does not help.

Terminal
docker run -d --gpus all -p 127.0.0.1:8888:8888 your-image

In Compose, write the same thing as ports: ["127.0.0.1:8888:8888"]. Docker releases older than 28.0.0 also let other hosts on the same network segment reach ports published to 127.0.0.1, so check docker version and keep authentication on the app regardless. From your laptop, ssh -N -L 8888:127.0.0.1:8888 quanta-gpu then opens the app at http://127.0.0.1:8888, as the SSH and VS Code guide explains.

Give PyTorch enough shared memory#

Docker gives a container 64 MB of /dev/shm unless you ask for more, and PyTorch's DataLoader workers and vLLM both use shared memory. When it runs out, training fails with a DataLoader error like this:

Output
ERROR: Unexpected bus error encountered in worker. This might be caused by insufficient shared memory (shm).

Add --ipc=host to share the host's IPC namespace, or set a size with --shm-size=16g. In Compose, the equivalents are ipc: host and shm_size: 16gb.

Common errors and their fixes#

Most GPU container failures print one of these:

ErrorWhat it meansFix
could not select device driver "" with capabilities: [[gpu]], or failed to discover GPU vendor from CDI from Docker 29.3 onDocker cannot find the NVIDIA Container Toolkit, or was not restarted after installing itInstall the toolkit, run nvidia-ctk runtime configure --runtime=docker and restart Docker
unknown or invalid runtime name: nvidia--runtime=nvidia was used before the runtime was registeredThe same fix
permission denied while trying to connect to the docker API, or the Docker daemon socket before Docker 29ubuntu is not in the docker groupsudo usermod -aG docker $USER, then log out and back in, or use sudo
docker.io/nvidia/cuda:latest: not found, or manifest for nvidia/cuda:latest not foundnvidia/cuda has no latest tagName a full tag, such as 12.8.1-base-ubuntu22.04
unsatisfied condition: cuda>=13.0The image was built for a newer CUDA than the driver supportsPick a tag from the image table above
The NVIDIA driver on your system is too old from PyTorch in a container that startedThe same mismatch, where Docker handed over the GPU through CDI without the checkPick a tag from the image table above
no kernel image is available for execution on the deviceThe image's CUDA or PyTorch build has no code for this GPU, such as a CUDA 12.6 build on BlackwellUse an image built with CUDA 12.8 or newer
Failed to initialize NVML: Unknown Error in a container that worked beforeThe container lost GPU access after a systemctl daemon-reload on the host, a known issue of the toolkit's hook pathDelete and recreate the container
Unexpected bus error encountered in worker/dev/shm is full at its 64 MB defaultAdd --ipc=host or --shm-size
executable file not found in $PATH for nvidia-smi, or nvidia-smi: command not found in a container shellThe container started without --gpusAdd --gpus all
no space left on device during a pullImages and models filled the VM's diskdocker system df shows the usage, and docker image prune -a removes unused images

Keep your setup in a script#

Every new instance starts from the same clean image, because stopping an instance terminates it and deletes its disk. The toolkit, the docker group change and every pulled image go with it. Keep the setup in a script on your laptop:

Terminal
#!/usr/bin/env bash
# gpu-docker-setup.sh: run on each new instance as ubuntu
set -euo pipefail

TOOLKIT_VERSION=1.20.1-1

if ! sudo docker run --rm --gpus all ubuntu nvidia-smi > /dev/null 2>&1; then
  sudo apt-get update
  sudo apt-get install -y --no-install-recommends ca-certificates curl gnupg2
  curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor --yes -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
  curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
    sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
    sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
  sudo apt-get update
  sudo apt-get install -y \
    nvidia-container-toolkit=$TOOLKIT_VERSION \
    nvidia-container-toolkit-base=$TOOLKIT_VERSION \
    libnvidia-container-tools=$TOOLKIT_VERSION \
    libnvidia-container1=$TOOLKIT_VERSION
  sudo nvidia-ctk runtime configure --runtime=docker
  sudo systemctl restart docker
fi

sudo usermod -aG docker "$USER"
sudo docker pull pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtime

Run it from your laptop with ssh quanta-gpu 'bash -s' < gpu-docker-setup.sh, and change the image tag to the one the image table gives for your driver. Copy results out of the container before you stop the instance: transferring files to and from a GPU server covers rsync, object storage and the Hugging Face Hub, and the guides index lists the other setups that run on the Bare Metal template.

Docker GPU questions#

Do I need to install CUDA on the host?

No. The host needs only the NVIDIA driver, which QuantaCloud's Bare Metal template already has. The CUDA runtime comes inside the image, and the toolkit mounts the driver into the container when it starts. Install the CUDA toolkit on the host only to compile CUDA code outside containers.

Do I need nvidia-docker?

No. The --gpus flag has been part of Docker since version 19.03, and the NVIDIA Container Toolkit is what makes it work. nvidia-docker was a wrapper script from the older nvidia-docker2 package.

Can I run my own Docker image on QuantaCloud?

Yes, on the Bare Metal template: pull it and run it with --gpus all. There is no custom-image template that starts your container for you at launch, and pulled images are deleted with the disk when you stop the instance. The GPU VPS page describes what the template includes.

Why does my container work with sudo but not without it?

ubuntu is not in the docker group yet, or the group change has not reached your session. Run sudo usermod -aG docker $USER, then log out and SSH back in, and docker ps works without sudo.


My order on a fresh instance never changes: nvidia-smi, then docker run --rm --gpus all ubuntu nvidia-smi, then an image tag chosen from the driver's CUDA version, with every port published on 127.0.0.1. Put the toolkit install and the image pulls in a script, because the next instance starts clean. Start on an RTX A6000 with the Bare Metal template, and move to the vLLM Docker guide when you are ready to serve a model.

Launch an Ubuntu GPU VM with Docker

Keep building

Choose your next step.