GPU guide

vLLM on multiple GPUs: tensor and pipeline parallelism

Split a model across 2, 4 or 8 GPUs with vLLM: tensor or pipeline parallelism, the flags, how to check the GPU links, and which VMs to rent.

Faiz Ahmed14 min read

The short answer is one flag: --tensor-parallel-size N splits every layer of a model across N GPUs in the same machine, and vLLM's docs recommend it when a model is too large for one GPU but fits on one node. Tensor parallelism exchanges data between the GPUs inside every layer, so it wants a fast link such as NVLink. On GPUs that talk over PCIe only, such as the L40S, vLLM's docs recommend --pipeline-parallel-size N instead, which gives each GPU a block of whole layers. QuantaCloud rents 2, 4 or 8 GPUs in one VM, and nvidia-smi topo -m shows which link your VM actually has.

Tensor, pipeline and data parallelism#

vLLM has four ways to use several GPUs, and they differ in what each GPU holds and how often the GPUs talk.

ModeFlagEach GPU holdsTraffic between GPUsUse it when
Tensor parallel--tensor-parallel-size NA slice of every layer, and its share of the KV cacheTwo all-reduces in every layer, on every forward passThe model does not fit one GPU and the GPUs share a fast link
Pipeline parallel--pipeline-parallel-size NWhole layers: on a 64-layer model at N=2, layers 1 to 32 on one GPU and 33 to 64 on the otherActivations handed from one stage to the nextThe GPUs talk over PCIe only, or the GPU count does not divide the model evenly
Data parallel--data-parallel-size NA full copy of the modelNone for a dense model: each copy serves its own requestsThe model fits one GPU and you want more requests per second
Expert parallel--enable-expert-parallelA share of a mixture-of-experts model's expertsAll-to-all exchanges so each token reaches its expertsLarge MoE models, together with tensor or data parallelism

The traffic column is the one that matters on rented hardware. The Megatron-LM authors, whose tensor-parallel algorithm vLLM implements, advise tensor parallelism only up to the GPUs in one server, because its all-reduce communication gets expensive over slower links, while pipeline parallelism uses cheaper point-to-point transfers. The InfiniBand vs NVLink guide explains the links themselves.

When one GPU is not enough#

The rule I follow is that the weights plus the KV cache you need must fit inside 92% of the GPUs you split across, because vLLM claims --gpu-memory-utilization 0.92 of every GPU by default. The weights are paid once, however many GPUs share them, so a second GPU adds far more room for the cache than the first one had. These figures are our calculation from the checkpoint sizes on Hugging Face and the memory each card reports in nvidia-smi, with ECC on for the 48 GB cards, ignoring activations:

Model and precisionWeightsGPUsvLLM budgetLeft for the KV cache
Qwen3-32B, BF1661.0 GiB1x 48 GB41.4 GiBDoes not fit
Qwen3-32B, BF1661.0 GiB2x 48 GB82.8 GiB21.8 GiB, about 89k tokens
Llama 3.3 70B, BF16131.4 GiB1x H200 NVL, 141 GB129.2 GiBDoes not fit, short by 2.2 GiB
Llama 3.3 70B, BF16131.4 GiB4x 48 GB165.6 GiB34.1 GiB, about 112k tokens
Llama 3.3 70B, BF16131.4 GiB2x H200 NVL258.3 GiB126.9 GiB, about 416k tokens
Llama 3.3 70B, BF16131.4 GiB8x A100 80GB588.8 GiB457.4 GiB, about 1.5M tokens
Llama 3.3 70B, FP867.7 GiB1x H200 NVL129.2 GiB61.5 GiB, about 200k tokens
Llama 3.3 70B, FP867.7 GiB2x H200 NVL258.3 GiB190.7 GiB, about 625k tokens

Qwen3-32B stores 256 KiB of KV cache per token and Llama 3.3 70B stores 320 KiB (2 x layers x 8 KV heads x 128 x 2 bytes), which turns the gigabytes into tokens. The FP8 rows use RedHatAI's per-channel checkpoint of Llama 3.3 70B, a 72.7 GB download, and the vLLM quantization guide covers which GPUs run FP8 at full speed. On 2x H200 NVL, the FP8 70B model gets three times the cache pool it has on one card, which is the real case for the second GPU: more long conversations at once, not just room for the weights. The KV cache guide has the formula, and how much VRAM you need covers other workloads.

If the model already fits one GPU, splitting it is the wrong tool for throughput. vLLM's own first rule is that distributed inference is probably unnecessary for a model that fits a single GPU, and --data-parallel-size 2 gives you two independent copies behind one endpoint instead.

Run vLLM across several GPUs#

The setup is the one in the vLLM Docker guide: the Bare Metal template, the driver check that picks v0.30.0 (R580 or newer) or v0.30.0-cu129, the GPU check and an API key in VLLM_API_KEY. On a 2-GPU VM, tensor parallelism is one extra flag:

Terminal
docker run -d --name vllm --gpus all --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v vllm-cache:/root/.cache/vllm \
  -e HF_TOKEN="${HF_TOKEN:-}" \
  "vllm/vllm-openai:$VLLM_TAG" \
  Qwen/Qwen3-32B \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --api-key "$VLLM_API_KEY"

For pipeline parallelism, swap the flag for --pipeline-parallel-size 2. For two copies of a model that fits one GPU, use --data-parallel-size 2, and on an 8-GPU VM --data-parallel-size 4 --tensor-parallel-size 2 runs four copies of a model split over two GPUs each. The modes also combine: vLLM's docs show --tensor-parallel-size 4 --pipeline-parallel-size 2 on eight GPUs, and the product of the sizes has to equal the GPUs you use.

--ipc=host is not optional here. vLLM's install docs say PyTorch shares memory between its processes, particularly for tensor parallel inference, and the flag gives the container the host's shared memory instead of Docker's small default. vLLM starts one worker process per GPU with Python multiprocessing, its default runtime on a single node, and each worker takes 92% of its own GPU.

Two log lines tell you whether the split worked. GPU KV cache size: N tokens is the cache pool shared by all requests, and Maximum concurrency for 32,768 tokens per request: X.XXx says how many full-length requests fit at once. Compare them with the same model on fewer GPUs.

The one thing I always check before choosing tensor parallelism is the link between the GPUs:

Terminal
nvidia-smi topo -m
nvidia-smi topo -p2p n
nvidia-smi topo -p2p p

The first command prints a matrix with one cell for every pair of GPUs. The other two ask the driver whether each pair can reach the other directly, over NVLink and over PCIe. NVIDIA's legend for the matrix:

EntryPath between the two GPUs
NV#A bonded set of # NVLinks
PIXA single PCIe switch
PXBSeveral PCIe switches, without the CPU's host bridge
PHBPCIe and a PCIe host bridge, typically the CPU
NODEPCIe and the link between host bridges inside one NUMA node
SYSPCIe and the link between NUMA nodes, usually between CPU sockets

vLLM runs the same check at startup. It uses NVLink detection to decide whether its own all-reduce kernel can run, and when it cannot, it logs one of two warnings and falls back to NCCL: Custom allreduce is disabled because it's not supported on more than two PCIe-only GPUs, or Custom allreduce is disabled because your platform lacks GPU P2P capability or P2P test failed. Neither stops the server, but on 4 or 8 GPUs the first one means vLLM did not find NVLink between every pair of GPUs.

NCCL makes the same choice underneath. Its peer-to-peer transport copies directly between GPUs over NVLink or PCIe, and when peer-to-peer cannot happen it falls back to shared host memory. NCCL_DEBUG=INFO in the container's environment turns on NCCL's own log for a closer look, and the DDP and FSDP guide runs the same checks before a training job.

QuantaCloud configurations with 2, 4 or 8 GPUs#

Every multi-GPU configuration is a single VM, and the GPU model decides what link is even possible. This is the 2026-09-27 catalog, with NVIDIA's specification for each card:

GPUMemory per GPUGPUs per VMNVLink on the cardCatalog flag
RTX A600048 GB1, 2, 42-way bridge, 112.5 GB/sNot stated
L4048 GB1, 2, 4NoneNot stated
L40S48 GB1, 2, 4, 8NoneNot stated
RTX 6000 Ada48 GB1, 2, 4, 8NoneNot stated
A100 80GB PCIe80 GB2, 42-GPU bridge, 600 GB/sPCIe
A100 80GB SXM480 GB1, 8600 GB/s through the HGX boardSXM and NVLink
H100 PCIe80 GB1, 22-card bridge, 600 GB/sPCIe
RTX PRO 6000 Blackwell96 GB1, 2, 4Not supportedNot stated
H200 NVL141 GB1, 22- or 4-way bridge, 900 GB/s per GPUNVLink

A card that supports a bridge does not mean the VM has one fitted, and the SXM, PCIe and NVLink guide explains how the form factors differ. The catalog flags NVLink on the A100 SXM4 and H200 NVL offers, and I treat that as unconfirmed until nvidia-smi topo -m on the VM shows NV entries. Live prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 6 Oct 2026, 00:22 UTC

Which mode to pick on each VM#

The decision follows the link. If topo -m shows NV entries between every pair, use tensor parallelism across all the GPUs, which is the case the A100 SXM4 and H200 NVL flags point to. If it shows PCIe paths, start with pipeline parallelism, as vLLM's docs advise for the L40S, and on the L40, RTX 6000 Ada and RTX PRO 6000 that is the only case there is. With two PCIe GPUs, I would still try tensor parallelism once and compare the output tokens per second, because two GPUs is the one PCIe case where vLLM's own all-reduce kernel can still run, provided nvidia-smi topo -p2p p shows peer-to-peer support between them.

For a model that fits one GPU, rent the GPU count you need for traffic and use data parallelism. I would run an 8B model on a 4x RTX A6000 VM as four copies, one per GPU, rather than one copy split four ways: no GPU waits for another, and each copy keeps a whole GPU of KV cache. VMs with 1 to 4 GPUs are running in about 3 minutes (median), and 8-GPU VMs in about 10.

Launch 2x RTX A6000 for vLLM

Mistakes on a multi-GPU server#

Five rules in vLLM's code and docs explain the usual surprises on several GPUs: the first three stop the server, and the last two cost cache.

SymptomCauseFix
Total number of attention heads (28) must be divisible by tensor parallel size (8)The tensor-parallel size has to divide the model's attention heads. DeepSeek-R1-Distill-Qwen-7B has 28Use a size that divides it (2 or 4 here), or pipeline parallelism
Pipeline parallelism is not supported for this modelThe model's vLLM implementation lacks pipeline supportUse tensor parallelism
A crash or hang while the workers start inside DockerThe container lacks shared memory for the worker processesAdd --ipc=host, then run vLLM's NCCL test script from its troubleshooting docs
Less KV cache than the arithmetic saysFewer KV heads than GPUs: Qwen3-235B-A22B has 4, so at TP=8 each one is stored twiceBudget for the duplicate, or keep the tensor-parallel size at or below the KV-head count when memory allows
A DeepSeek MLA model gains little cache per GPUTensor parallelism keeps a full copy of the compressed MLA cache on every GPUvLLM suggests data parallelism for the attention layers of MLA models

NCCL_P2P_DISABLE=1 is the documented escape hatch when the NCCL test hangs, and vLLM's troubleshooting page says to use it only as a temporary workaround because it can cost performance: NCCL then moves the data through host memory. Uneven GPU counts have their own answer: pipeline parallelism splits along layers and accepts uneven splits, so vLLM's docs suggest it when the GPU count does not divide the model.

A model too large for the GPUs you give it stops at startup with one of vLLM's out-of-memory errors, and the vLLM out of memory guide goes through each of them.

Big models also mean big downloads. Llama 3.3 70B in BF16 is 141.1 GB, and stopping a QuantaCloud instance deletes its disk, so every new launch downloads the weights again. Meta's repository is gated, so it needs an HF_TOKEN with access, while RedHatAI's FP8 build is not. Keep the launch command in a script, pre-fetch with the hf CLI as the Hugging Face download guide shows, and expect vllm serve to spend its first minutes loading.

vLLM multi-GPU FAQ#

No. Tensor, pipeline and data parallelism all run over PCIe. NVLink makes tensor parallelism cheaper, because tensor parallelism runs two all-reduces in every layer, which is why vLLM's docs recommend pipeline parallelism on GPUs without it, such as the L40S.

How many GPUs should I split a model across?

The fewest that hold the weights plus the KV cache you need at 92% utilization. Past that, add copies with --data-parallel-size for more requests per second, and add tensor-parallel GPUs only when you need lower latency per request or more cache per request.

Can one vLLM server span two QuantaCloud VMs?

Not on demand. On-demand instances are single VMs with up to 8 GPUs, and vLLM's multi-node setup (tensor parallelism inside each node, pipeline parallelism across nodes, over a private network because the traffic is not encrypted) needs servers built for it. We build multi-node clusters with InfiniBand to order, as GPU clusters describes: send a capacity brief and the configuration, lead time and terms come back in writing.

Does tensor parallelism make a small model faster?

Not reliably. Every layer then waits for two all-reduces between the GPUs, and a small model spends a larger share of each step on that exchange. If the model fits one GPU, run it on one GPU, or run one copy per GPU with data parallelism.


My rule for vLLM on several GPUs: fit the model on the fewest GPUs that leave room for your KV cache, use tensor parallelism when nvidia-smi topo -m shows NVLink, pipeline parallelism when it shows PCIe, and data parallelism once the model fits one GPU. Start on a 2x RTX A6000 VM for 30B-class models, and rent the 8x A100 SXM4 when a 70B model in BF16 needs room to serve many users. Self-hosted LLM inference covers the rest of the serving setup.

Launch an 8x A100 SXM4 VM

Keep building

Choose your next step.