GPU guide

Install vLLM and serve a model with vllm serve

Install vLLM 0.30.0 with uv or pip, match the wheel to your NVIDIA driver (CUDA 13 or cu129), then serve an OpenAI-compatible API with vllm serve.

Faiz Ahmed8 min read

The fastest way to install vLLM is uv in a fresh virtual environment, and the one check that decides the command is your NVIDIA driver version. vLLM 0.30.0 on PyPI is a CUDA 13.0 build that needs an R580 or newer driver. On an older driver you install the cu129 wheel from vLLM's GitHub release instead. After that, vllm serve gives you an OpenAI-compatible API on port 8000.

If you would rather run a container than manage a Python environment, the vLLM Docker guide uses the official image with the same flags. To weigh vLLM against other servers first, see vLLM vs Ollama and vLLM vs SGLang.

Launch a GPU VM#

Use the Bare Metal template: Ubuntu 22.04 with the NVIDIA driver and Docker. The PyTorch + Jupyter template runs its PyTorch inside a container, and vLLM pins its own torch==2.13.0, so installing vLLM into that environment invites version fights. A single RTX A6000 with 48 GB is enough for an 8B model in BF16.

Launch an RTX A6000 (Bare Metal)

Most single-GPU VMs are running in about 3 minutes (median). Connect as ubuntu (connecting over SSH) and look at the GPU and driver:

Terminal
ssh ubuntu@YOUR_INSTANCE_IP
nvidia-smi --query-gpu=name,driver_version,compute_cap --format=csv

Match the wheel to your driver#

The rule I follow is to pick the wheel from the driver branch, never from the CUDA Toolkit version, because the wheels bring their own CUDA libraries and only need a new enough driver.

DriverInstallWhy
580 or newervllm from PyPIThe default wheel is a CUDA 13.0 build with torch 2.13.0, and CUDA 13 needs R580
Older than 580vllm-0.30.0+cu129 from the GitHub releaseThe CUDA 12.9 build runs on older drivers

The GPU generation rarely decides the wheel. vLLM needs compute capability 7.5 or newer, and both wheels include kernels for every GPU QuantaCloud offers on demand. The cu129 wheels skip only compute capability 10.3 (B300, GB300) and 12.1 (GB10).

GPUArchitectureCompute capabilityNote
A100 80 GBAmpere8.0No FP8 math
RTX A6000Ampere8.6No FP8 math
L40S, L40, RTX 6000 AdaAda Lovelace8.9FP8
H100 PCIe, H200 NVLHopper9.0FP8
RTX PRO 6000 BlackwellBlackwell12.0FP8 and FP4, needs an R570 or newer driver

nvcc --version does not matter here, because the wheels bring their own CUDA libraries. Checking your GPU, driver and CUDA version explains the three CUDA numbers you will meet.

Install uv and create an environment#

uv installs in one line, and it can download a Python build for you. That matters on Ubuntu 22.04, whose own Python is 3.10.

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env

mkdir -p ~/vllm && cd ~/vllm
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate

vLLM's docs recommend a fresh environment for a reason: its compiled kernels are tied to one PyTorch build. Do not install vLLM next to an existing PyTorch, and do not upgrade torch inside this environment afterwards.

Install vLLM#

On an R580 or newer driver, install from PyPI. --torch-backend=auto lets uv read the driver and pick the matching PyTorch index.

Terminal
uv pip install vllm==0.30.0 --torch-backend=auto

On an older driver, install the cu129 wheel from the release, with PyTorch from its cu129 index:

Terminal
export VLLM_VERSION=0.30.0
export CUDA_VERSION=129
uv pip install "https://github.com/vllm-project/vllm/releases/download/v${VLLM_VERSION}/vllm-${VLLM_VERSION}+cu${CUDA_VERSION}-cp38-abi3-manylinux_2_28_$(uname -m).whl" \
  --extra-index-url "https://download.pytorch.org/whl/cu${CUDA_VERSION}"

Plain pip works too. pip install vllm==0.30.0 inside the environment installs the same CUDA 13.0 build, and the cu129 command above works with pip install in place of uv pip install. I use uv because it resolves and downloads faster.

Check that the install works#

Three numbers tell you the install is sound: the vLLM version, the CUDA version of the PyTorch build, and whether PyTorch can see the GPU.

Terminal
vllm --version
python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available()); print(torch.cuda.get_device_name(0), torch.cuda.get_arch_list())"

torch.version.cuda should print 12.9 for the cu129 wheel and a 13.x version for the default one: 13.0 with plain pip and 13.2 with uv, because --torch-backend=auto picks the newest PyTorch CUDA build that an R580 or newer driver allows. torch.cuda.is_available() must print True. If it prints False with a warning that the NVIDIA driver is too old, you installed the CUDA 13 build on a pre-580 driver: delete .venv and use the cu129 command. The arch list should contain your GPU's compute capability or a lower one with the same major number: sm_86 for an RTX A6000 (it also covers the Ada cards) or sm_120 for an RTX PRO 6000.

Serve a model with vllm serve#

vllm serve downloads the model from Hugging Face, loads it and starts the API server. Bind it to 127.0.0.1 and set an API key, because by default it listens on every interface of a VM that has a public IP.

Terminal
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"

vllm serve Qwen/Qwen3-8B \
  --host 127.0.0.1 --port 8000 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --api-key "$VLLM_API_KEY"

The first start downloads 16.4 GB of weights into ~/.cache/huggingface and then loads them. Wait for Application startup complete.

A few defaults are worth knowing. The port is 8000. vLLM takes 92% of GPU memory (--gpu-memory-utilization 0.92) for weights, activations and the KV cache. The served model name is the Hugging Face id unless you pass --served-model-name. Sampling defaults come from the model's generation_config.json unless you pass --generation-config vllm. --reasoning-parser qwen3 returns Qwen3's thinking separately from its answer. --max-model-len caps the context, which is the main lever on memory, and the KV cache guide shows how to size it. Write it out in full: vLLM reads 32k as 32,000 tokens and 32K as 32,768.

Call the OpenAI-compatible API#

vLLM serves the OpenAI API shapes: /v1/chat/completions, /v1/completions, /v1/responses, /v1/embeddings and /v1/models, plus an Anthropic-style /v1/messages. Test it from a second SSH session. A new session has neither the key nor the environment, so set both first:

Terminal
export VLLM_API_KEY=PASTE_THE_KEY_YOU_PRINTED
source ~/vllm/.venv/bin/activate

curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "What does vllm serve do?"}], "max_tokens": 1000}'

Qwen3 thinks before it answers, and the thinking counts against max_tokens, so leave it room or the answer comes back empty. The openai Python package is already in the environment, because vLLM depends on it:

Python
import os
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key=os.environ["VLLM_API_KEY"])
reply = client.chat.completions.create(
    model="Qwen/Qwen3-8B",
    messages=[{"role": "user", "content": "What does vllm serve do?"}],
    max_tokens=1000,
)
print(reply.choices[0].message.content)

From your laptop, forward the port over SSH, then run the same code with pip install openai and the key in VLLM_API_KEY:

Terminal
ssh -N -L 8000:127.0.0.1:8000 ubuntu@YOUR_INSTANCE_IP

The API key guards only the /v1-style routes, so keep the server on 127.0.0.1 and use the tunnel. The Docker guide's section on exposing vLLM shows a reverse proxy with TLS for when an app on another machine needs it, and the SSH and VS Code guide covers the tunnel in more depth.

Run batch inference without a server#

vLLM does not need the HTTP server at all when the prompts come from your own script. The LLM class runs the same engine offline, and llm.chat applies the model's chat template for you. Stop vllm serve first, because both claim 92% of the GPU.

Python
from vllm import LLM, SamplingParams

llm = LLM(model="Qwen/Qwen3-8B", max_model_len=32768)
params = SamplingParams(temperature=0.7, max_tokens=1000)
conversations = [
    [{"role": "user", "content": "Summarize what a KV cache is."}],
    [{"role": "user", "content": "Name three uses for a 48 GB GPU."}],
]
for output in llm.chat(conversations, params):
    print(output.outputs[0].text)

I would use it for evaluation runs and dataset generation, where a server is only overhead. Anything another program has to call belongs behind vllm serve.

Keep it running after you log out#

vllm serve stops when your SSH session ends. The simplest fix is tmux:

Terminal
sudo apt-get update && sudo apt-get install -y tmux
tmux new -s vllm

Activate the environment and start vllm serve inside the tmux session, then press Ctrl+B followed by D to detach. tmux attach -t vllm brings it back. This survives a dropped connection, not a stopped instance: stopping a QuantaCloud instance terminates it and deletes its disk, including the environment and the model cache, so the install and the first download happen again on the next launch.

What vLLM will and will not load#

vLLM reads the quantization format from the checkpoint's quantization_config, so a pre-quantized model needs no extra flag. AWQ, GPTQ, FP8, MXFP4, NVFP4 and compressed-tensors are built in. FP8 math needs compute capability 8.9 or newer, and the Ampere cards run FP8 checkpoints weight-only. MXFP4, the format of gpt-oss, needs 8.0 or newer (gpt-oss GPU requirements). GPTQ checkpoints quantized with group activation ordering (desc_act=True) no longer load in 0.30.0, so pick one made with desc_act=False.

bitsandbytes and GGUF moved out of vLLM into separate plugins, vllm-bnb-plugin and vllm-gguf-plugin, and vLLM describes its GGUF support as highly experimental. If your model only exists as GGUF, Ollama is the natural server for it (the Ollama API on a remote GPU). llama.cpp runs GGUF files as well.


My rule is pip or uv when vLLM is one part of a Python project on the VM, and Docker when vLLM is the service. Either way, check the driver before you install, keep the server on 127.0.0.1, and size the GPU from the weights plus the KV cache you need. If an 8B model at 32k context is your target, an RTX A6000 is the place to start.

Launch a GPU for vLLM

Keep building

Choose your next step.