The fastest way to run vLLM with Docker is the official vllm/vllm-openai image from Docker Hub, pinned to a release tag, with port 8000 published on 127.0.0.1 only and an API key set. On QuantaCloud that means launching the Bare Metal template (Ubuntu 22.04 with the NVIDIA driver and Docker), checking the driver, and running one container. QuantaCloud has no vLLM template and no managed endpoint, so this is a do-it-yourself deployment: the VM, the container and their security are yours to run.
What you end up with is an OpenAI-compatible server (/v1/chat/completions, /v1/completions, /v1/models and the rest) that any OpenAI SDK can call by changing base_url.
Choose a GPU for your model#
The rule I follow is that the weights plus the KV cache for your context have to fit inside 92% of the card, because that is how much memory vLLM claims by default. The weights are easy to read off the model repo. The KV cache is the part people forget: it grows with every token of context and every concurrent request, and the KV cache guide shows how to size it.
| Model | Precision | Weights | GPU that fits |
|---|---|---|---|
| Qwen3-8B | BF16 | 16.4 GB | RTX A6000, 48 GB |
| Qwen3-32B | BF16 | 65.5 GB | A100 or H100 PCIe, 80 GB (tight), RTX PRO 6000, 96 GB, or two 48 GB cards with pipeline parallelism |
| Qwen3-32B | FP8 | 34.3 GB | L40S, L40 or RTX 6000 Ada, 48 GB (borderline at 32k) |
| gpt-oss-120b | MXFP4 | 65.2 GB | H100 PCIe, 80 GB (tight), RTX PRO 6000 or H200 NVL |
| Llama 3.3 70B | FP8 | 72.7 GB | H200 NVL, 141 GB, or RTX PRO 6000 at a shorter context |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Two hardware rules decide the precision column. FP8 math needs compute capability 8.9 or newer: Ada (L40S, L40, RTX 6000 Ada), Hopper (H100, H200) and Blackwell (RTX PRO 6000). On the Ampere cards, the RTX A6000 and A100, vLLM runs FP8 checkpoints weight-only through its Marlin kernels, so you get the memory saving but not the FP8 speed. The Ada cards get FP8 math only from per-channel checkpoints such as RedHatAI/Qwen3-32B-FP8-dynamic: vLLM 0.30.0 runs block-scaled ones, including Qwen's own Qwen3-32B-FP8, weight-only below Hopper. MXFP4, the format gpt-oss ships in, needs compute capability 8.0 or newer. The gpt-oss sizing guide covers that model in detail, and how much VRAM you need covers the rest.
Launch an Ubuntu GPU VM#
Launch the Bare Metal template on the GPU you picked (the deploy flow). This walkthrough uses one RTX A6000 and Qwen3-8B.
Launch an RTX A6000 for vLLMMost single-GPU VMs are running in about 3 minutes (median). The instance has a public IP, and you connect as ubuntu with the SSH key you picked at deploy time (connecting over SSH):
ssh ubuntu@YOUR_INSTANCE_IP
The offer card shows the included disk: 256 GB on a single RTX A6000 in the 2026-09-27 catalog, which easily holds the image (an 8.7 GB download, larger once unpacked) and the 16.4 GB of weights. That disk is deleted when you stop the instance, so everything below runs again on the next launch.
Check the driver first#
The one thing I always check before pulling the image is the driver version. The default vLLM 0.30.0 image, v0.30.0 (the same image as latest), is a CUDA 13.0 build and needs an R580 or newer driver. The v0.30.0-cu129 tag is the CUDA 12.9 build for older drivers.
nvidia-smi --query-gpu=name,driver_version,memory.total,compute_cap --format=csv
| Driver | Image tag | Download (amd64, compressed) |
|---|---|---|
| 580 or newer | vllm/vllm-openai:v0.30.0 | 8.7 GB |
| Older than 580 | vllm/vllm-openai:v0.30.0-cu129 | 13.7 GB |
Set the tag once, so every later command uses the right one:
DRIVER_MAJOR=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader | head -n1 | cut -d. -f1)
if [ "$DRIVER_MAJOR" -ge 580 ]; then VLLM_TAG=v0.30.0; else VLLM_TAG=v0.30.0-cu129; fi
echo "vllm/vllm-openai:$VLLM_TAG"
The CUDA 13 image also has a compatibility mode for R535 and R570 drivers on select datacenter and professional GPUs (VLLM_ENABLE_CUDA_COMPATIBILITY=1), but the cu129 tag is the simpler fix. Checking your GPU, driver and CUDA version explains the driver and CUDA numbers that nvidia-smi prints.
Check that Docker can see the GPU#
Docker needs the NVIDIA Container Toolkit to hand the GPU to a container. First make sure ubuntu can use Docker without sudo:
docker ps
If that fails with permission denied on the Docker socket, add yourself to the docker group, then log out and SSH back in. Membership of that group is equivalent to root, which is fine on a single-user VM. The new session starts without VLLM_TAG, so run the tag snippet from the previous section again.
sudo usermod -aG docker $USER
Then run a quick GPU check:
docker run --rm --gpus all ubuntu nvidia-smi
If it prints the same GPU table as the host, you are done. If Docker answers could not select device driver "" with capabilities: [[gpu]], install and configure the toolkit:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Run vLLM#
Pin the tag rather than latest, which moves to each new release, publish the port on 127.0.0.1 only, and set an API key. Generate the key first and keep it, because every client sends it as a Bearer token.
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v vllm-cache:/root/.cache/vllm \
-e HF_TOKEN="${HF_TOKEN:-}" \
"vllm/vllm-openai:$VLLM_TAG" \
Qwen/Qwen3-8B \
--max-model-len 32768 \
--reasoning-parser qwen3 \
--api-key "$VLLM_API_KEY"
docker logs -f vllm
The image's entrypoint is vllm serve, so everything after the image name is a vllm serve argument: the model, then the flags. --ipc=host gives PyTorch the shared memory it needs. The Hugging Face mount keeps the weights on the VM's disk when you recreate the container, and the vllm-cache volume keeps the torch.compile artifacts so the next start compiles less. HF_TOKEN only matters for gated repos such as Llama. Qwen3-8B is Apache-2.0 and not gated. --reasoning-parser qwen3 returns Qwen3's thinking separately from its answer.
I set --max-model-len explicitly and leave --gpu-memory-utilization at its default of 0.92, so the memory budget is predictable. Write the length out in full: vLLM reads 32k as 32,000 tokens and 32K as 32,768. Wait for Application startup complete in the log. The container runs as root by default, and the image also supports its built-in vllm user (--user 2000:0) with the caches mounted under /home/vllm instead.
You only need a Dockerfile of your own to add packages, and it should start from the same tag:
FROM vllm/vllm-openai:v0.30.0
RUN uv pip install --system "vllm[audio]==0.30.0"
The same deployment with Docker Compose#
Compose keeps the flags in a file and restarts the server if it crashes. Check the plugin with docker compose version, then write the secrets to ~/vllm/.env and save the Compose file below next to it as compose.yaml:
mkdir -p ~/vllm && cd ~/vllm
cat > .env <<EOF
VLLM_TAG=$VLLM_TAG
VLLM_API_KEY=$VLLM_API_KEY
HF_TOKEN=${HF_TOKEN:-}
EOF
chmod 600 .env
services:
vllm:
image: vllm/vllm-openai:${VLLM_TAG}
command: ["Qwen/Qwen3-8B", "--max-model-len", "32768", "--reasoning-parser", "qwen3"]
environment:
- VLLM_API_KEY=${VLLM_API_KEY}
- HF_TOKEN=${HF_TOKEN}
ports:
- "127.0.0.1:8000:8000"
ipc: host
volumes:
- ${HOME}/.cache/huggingface:/root/.cache/huggingface
- vllm-cache:/root/.cache/vllm
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 5s
retries: 5
start_period: 600s
volumes:
vllm-cache:
vLLM reads VLLM_API_KEY from the environment, which does the same job as --api-key. The health check can call /health without the key, and start_period gives the first start 10 minutes to download and compile before failures count. Remove the container from the previous section, then start the Compose version:
docker rm -f vllm
docker compose up -d
docker compose logs -f
Test the API#
Both calls need the key. From the VM:
curl -s http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $VLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "Explain the KV cache in two sentences."}], "max_tokens": 1000}'
Qwen3 thinks before it answers, and the thinking counts against max_tokens, so leave it room or the answer comes back empty.
From your laptop, open the SSH tunnel from the next section, set VLLM_API_KEY to the key you printed earlier, and point the OpenAI Python SDK (pip install openai) at the tunnel with base_url:
import os
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key=os.environ["VLLM_API_KEY"])
reply = client.chat.completions.create(
model="Qwen/Qwen3-8B",
messages=[{"role": "user", "content": "Explain the KV cache in two sentences."}],
max_tokens=1000,
)
print(reply.choices[0].message.content)
Expose it safely#
The default I recommend is an SSH tunnel, because then the API is reachable only through SSH on port 22. Run this on your laptop and leave it open:
ssh -N -L 8000:127.0.0.1:8000 ubuntu@YOUR_INSTANCE_IP
Your laptop's http://127.0.0.1:8000/v1 now reaches the server. The SSH and VS Code guide covers keys and port forwarding in more depth.
An API key alone is not enough to open port 8000 to the internet. --api-key protects only the /v1, /v2, /inference and /cohere routes. /invocations runs the same inference with no key at all, and /pause and /abort_requests let anyone pause generation or kill the requests in flight. Docker makes it worse: it routes published ports before ufw sees them, so -p 8000:8000 on all interfaces ignores your firewall. On Docker releases older than 28.0.0, hosts on the same network segment can even reach ports published to 127.0.0.1, so check docker version and keep the key on regardless.
When an app on another machine needs the API, put a reverse proxy with TLS in front and let only /v1 through, with the API key still on. Caddy fetches the certificate itself, which needs inbound ports 80 and 443 to reach the VM. Point a DNS record at the instance IP (again after every new launch), install Caddy from its apt repository, and replace /etc/caddy/Caddyfile with:
api.example.com {
handle /v1/* {
reverse_proxy 127.0.0.1:8000
}
handle {
respond 404
}
}
Then reload Caddy and open only SSH, HTTP and HTTPS:
sudo systemctl reload caddy
sudo ufw allow 22/tcp && sudo ufw allow 80/tcp && sudo ufw allow 443/tcp
sudo ufw --force enable
Tune memory and context#
The number to watch is the line vLLM prints once the model is loaded:
GPU KV cache size: N tokens, Maximum concurrency for 32,768 tokens per request: X.XXx
Qwen3-8B stores 144 KiB of KV cache per token (2 x 36 layers x 8 KV heads x 128 x 2 bytes), so one 32,768-token request needs 4.5 GiB. An RTX A6000 with ECC on reports 46,068 MiB, and 92% of that is about 41.4 GiB. Take away 15.3 GiB of weights and at most 26.1 GiB is left for the cache and activations, room for nearly six full-length requests at once (our calculation). With ECC off the card reports 49,140 MiB, 2.8 GiB more, and the memory.total column of the driver check shows which you have.
If the model does not fit, vLLM refuses to start with an error that begins To serve at least one request with the model's max seq len and names the largest length that would fit. Lower --max-model-len, or pass --max-model-len auto and let vLLM pick. --kv-cache-dtype fp8 halves the cache, but without calibration it uses scales of 1.0, so check the output quality on your own prompts. Prefix caching is on by default, so a long system prompt that repeats is stored once. The KV cache guide has the formula for any model, and fixing CUDA out of memory covers the other errors.
Use more than one GPU#
The flag for a multi-GPU VM is --tensor-parallel-size: add --tensor-parallel-size 2 after the model name on a 2x VM and vLLM splits every layer across both GPUs. --ipc=host is already set, and it matters here. Tensor parallelism talks between GPUs on every layer, so vLLM's docs recommend pipeline parallelism (--pipeline-parallel-size 2) on GPUs without NVLink, and the L40S, L40, RTX 6000 Ada and RTX PRO 6000 have none. The catalog flags NVLink on the A100 SXM4 and H200 NVL offers. Check what your VM actually has with nvidia-smi topo -m before you rely on it.
Benchmark method#
The method is the same on every GPU, so the results compare across GPUs. vLLM vs SGLang compares the two servers feature by feature.
- Server: the pinned image, the flags shown on the page,
--gpu-memory-utilizationat 0.92, and the weights already in the local cache. - Load time: seconds from
docker runto the first 200 from/healthwith cached weights. The first start, including the download, is timed separately. - Load:
vllm bench serveinside the same container, random prompts of 1,024 input and 256 output tokens with--ignore-eos, 32 prompts at concurrency 1 and 256 prompts at concurrency 16 (and at 64 where a page lists it), with a new--seedfor every run so the prefix cache cannot inflate the result. - Reported: output tokens per second, median time to first token, median time per output token, the KV cache line from the log, and the GPU, driver and peak memory from
nvidia-smi.
docker exec -e OPENAI_API_KEY="$VLLM_API_KEY" vllm vllm bench serve \
--backend vllm --model Qwen/Qwen3-8B \
--dataset-name random --random-input-len 1024 --random-output-len 256 --ignore-eos \
--max-concurrency 16 --num-prompts 256 --seed 16
The benchmark client reads the key from OPENAI_API_KEY. With Compose, use docker compose exec and the service name vllm instead.
Before you stop#
Stopping a QuantaCloud instance terminates it and deletes its disk, including the Hugging Face cache and the pulled image. The next launch downloads both again: 8.7 GB of image (13.7 GB for the cu129 tag) plus 16.4 GB for Qwen3-8B, or 65.2 GB for gpt-oss-120b. Keep the commands from this page in a script on your laptop and run it over SSH on each new instance, and see downloading Hugging Face models fast to shorten the wait. The first hour is charged at launch and the unused seconds are refunded when you stop (how billing works).
My rule for the first deployment is to start on an RTX A6000 with an 8B model and exactly these commands, then read the KV cache line. Move up to the RTX PRO 6000 or the H200 NVL only when that line says the context or concurrency you need does not fit. If you would rather manage a Python environment than a container, install vLLM with pip or uv instead.
Launch an Ubuntu GPU VM for vLLM