PCIe and SXM are two ways to mount a GPU in a server, and NVLink is a direct link between GPUs. A PCIe card plugs into a standard slot and talks to other GPUs over PCIe, 64 GB/s on a Gen4 x16 slot and 128 GB/s on Gen5, unless an NVLink bridge joins it to its neighbors. An SXM module sits on an NVIDIA HGX board, and on an 8-GPU board NVSwitch chips connect all eight over NVLink: 600 GB/s per GPU on the A100, 900 GB/s on the H100 and H200, and 1.8 TB/s on the B200 and B300. For a job on one GPU, none of this matters. For one model split across several GPUs, the link can decide the speed.
On QuantaCloud's on-demand VMs, the A100 comes in both versions, SXM4 and PCIe, and the H200 NVL is a PCIe card built for NVLink bridges. The L40S, L40, RTX 6000 Ada and RTX PRO 6000 Blackwell have no NVLink at all. Which links a VM actually exposes is something to check on the VM, and the last sections show how.
Launch an 8x A100 SXM4 VMPCIe, SXM, NVLink and NVSwitch#
These terms describe two different things: how a GPU is packaged, and how it talks to other GPUs.
| Term | What it is | Where you meet it |
|---|---|---|
| PCIe card | A GPU on a standard add-in board, in a PCIe slot | RTX A6000, L40S, H100 PCIe, H200 NVL, RTX PRO 6000 Blackwell |
| SXM module | A GPU module that mounts flat on an NVIDIA HGX baseboard instead of a slot | A100 SXM4, H100 SXM5, H200 SXM, B200, B300 |
| NVLink | NVIDIA's direct GPU-to-GPU link, built from several bonded links per GPU | Between SXM modules, or between PCIe cards joined by a bridge |
| NVLink bridge | A connector across the tops of 2 or 4 PCIe cards | RTX A6000 pairs, A100 PCIe and H100 PCIe pairs, H200 NVL groups of 2 or 4 |
| NVSwitch | Switch chips on the HGX board that connect every GPU to every other at full NVLink speed | 8-GPU HGX boards, and NVL72 racks with 72 GPUs |
An NVLink bridge connects only the cards it spans, so a bridged pair is fast between its two GPUs and back on PCIe for everything else. NVSwitch is what makes all eight GPUs on an HGX board equal: NVIDIA quotes 900 GB/s between any pair on an 8-GPU HGX H100 board.
NVLink and PCIe bandwidth#
NVIDIA quotes both links in gigabytes per second with the two directions added together. These are its figures:
| Link | Bandwidth per GPU | Used by |
|---|---|---|
| PCIe Gen4 x16 | 64 GB/s | RTX A6000, L40, L40S, RTX 6000 Ada, A100 |
| PCIe Gen5 x16 | 128 GB/s | H100 PCIe, H200 NVL, RTX PRO 6000 Blackwell |
| PCIe Gen6 x16 | 256 GB/s | HGX B300 |
| NVLink bridge, RTX A6000 | 112.5 GB/s | 2 cards |
| NVLink bridge, A100 PCIe and H100 PCIe | 600 GB/s | 2 cards |
| NVLink bridge, H200 NVL | 900 GB/s | 2 or 4 cards |
| NVLink, third generation (A100 SXM4) | 600 GB/s, 12 links | 8 GPUs on HGX A100 |
| NVLink, fourth generation (H100, H200) | 900 GB/s, 18 links | 8 GPUs through NVSwitch |
| NVLink, fifth generation (B200, B300, GB200, GB300) | 1,800 GB/s, 18 links | 8 GPUs on HGX, or 72 in an NVL72 rack |
The gap depends on the generation. NVIDIA puts the H100's NVLink at 7 times PCIe Gen5, the A100's 600 GB/s is about 9 times its PCIe Gen4 slot, and the RTX A6000's bridge is less than twice a Gen4 slot (our calculation: 900 / 128, 600 / 64 and 112.5 / 64). NVIDIA's NVLink page lists a sixth generation at 3,000 GB/s per GPU for the Vera Rubin platform, as a preliminary figure.
Between servers the link is the network, not NVLink or PCIe. InfiniBand vs NVLink covers that half, with NVIDIA's port speeds and the ratio between the two.
SXM vs PCIe for a single GPU#
The rule I follow for a single-GPU job is to ignore the link and compare memory, bandwidth and price. On one GPU there is no other GPU to talk to, so SXM only matters for what else it changes: a higher power limit, and on the H100 a different memory type.
| A100 SXM4 vs A100 PCIe | H100 SXM5 vs H100 PCIe | H200 SXM vs H200 NVL | |
|---|---|---|---|
| Memory | 80 GB HBM2e on both | 80 GB HBM3 vs 80 GB HBM2e | 141 GB HBM3e on both |
| Memory bandwidth | 2,039 vs 1,935 GB/s | 3.35 vs 2.0 TB/s | 4.8 TB/s on both |
| Tensor throughput | The same peak TFLOPS | 1,979 vs 1,513 dense FP8 TFLOPS | 3,958 vs 3,341 FP8 TFLOPS with sparsity |
| Max power | 400 W vs 300 W | Up to 700 W vs 350 W | Up to 700 W vs up to 600 W |
| On QuantaCloud | Both on demand | PCIe on demand, SXM built to order | NVL on demand, SXM built to order |
The H100 is the case where SXM changes the chip's speed as well: on a basket of 10 data analytics, AI and HPC applications, NVIDIA's own whitepaper puts the H100 PCIe at 65 percent of the SXM5's performance, at half the power. For the A100, the two versions have the same published TFLOPS and nearly the same bandwidth, so for one GPU I take whichever is cheaper when I check. The A100, H100 and H200 pages have the full specs and live prices, and the GPU comparisons set other pairs of cards side by side.
When the link matters#
The link matters in proportion to how much data your GPUs exchange, and that depends on how the work is split:
| How the work is split | What crosses between GPUs | How much the link matters |
|---|---|---|
| One job or one model per GPU | Nothing | Not at all |
| Data parallel training (DDP) | Every gradient, once per step | Some, more for larger models |
| Sharded training (FSDP, ZeRO stage 3) | Weights and gradients every step: the ZeRO paper counts 1.5 times the data of plain data parallel for stage 3 | More |
| Pipeline parallelism | Activations between stages, point to point | Least of the ways to split one model |
| Tensor parallelism | Partial results inside every layer | Most: keep it inside one NVLink domain |
The Megatron-LM authors turned this into a rule that still holds: tensor parallelism up to the GPUs in one server, pipeline parallelism across servers. vLLM applies the same logic inside one machine. Its docs recommend tensor parallelism across the GPUs of a node, and pipeline parallelism instead when the GPUs have no NVLink, naming the L40S as the example. On a 4x L40S or 4x RTX 6000 Ada VM, that means --pipeline-parallel-size 4 rather than --tensor-parallel-size 4, and on a VM where nvidia-smi topo -m shows NVLink, tensor parallelism is the usual choice. Tensor parallelism in vLLM walks through both, and for training, DDP and FSDP on one node and DeepSpeed ZeRO cover the data-parallel side.
What QuantaCloud's multi-GPU VMs are built from#
Every on-demand configuration is one VM with 1, 2, 4 or 8 GPUs, and the GPU model decides which links are possible:
| Configuration | What NVIDIA's GPU supports between cards | How the offers API labels it |
|---|---|---|
| 8x A100 SXM4 | NVLink at 600 GB/s through the HGX A100 board | SXM with NVLink |
| 2x or 4x A100 PCIe | An NVLink bridge for 2 GPUs, 600 GB/s | PCIe |
| 2x H100 PCIe | An NVLink bridge to one adjacent card, 600 GB/s | PCIe |
| 2x H200 NVL | A 2- or 4-way NVLink bridge, 900 GB/s per GPU | NVLink |
| 2x or 4x RTX A6000 | A 2-way NVLink bridge, 112.5 GB/s | No label |
| Up to 8x L40S or RTX 6000 Ada, up to 4x L40 or RTX PRO 6000 Blackwell | No NVLink: PCIe only | No label |
A label in the API is not proof that the VM exposes the link, and a card that supports a bridge does not always have one fitted. My rule is to read nvidia-smi topo -m on the VM before choosing a tensor-parallel size.
How to check the links on your VM#
The one thing I always run first on a multi-GPU VM is the topology matrix. Connect over SSH (the SSH guide has the setup) and run:
nvidia-smi topo -m
Each cell describes the path between two GPUs:
| Entry | Meaning |
|---|---|
| NV followed by a number | The GPUs share that many bonded NVLinks |
| PIX | The path crosses a single PCIe switch |
| PXB | The path crosses several PCIe switches, but not the CPU's host bridge |
| PHB | The path goes through a PCIe host bridge, usually the CPU |
| NODE | The path crosses host bridges inside one NUMA node |
| SYS | The path crosses the link between NUMA nodes, usually between CPU sockets |
If the matrix shows NV entries, check that the links are up:
nvidia-smi nvlink -s
For each active link it prints the bandwidth, and a link that is present but not active shows as Inactive. To see the PCIe side, query the link generation and width of every GPU:
nvidia-smi --query-gpu=index,pcie.link.gen.current,pcie.link.gen.max,pcie.link.width.current,pcie.link.width.max --format=csv
The current values can drop while a GPU is idle, so read them while a job is running. The number that settles the question is NCCL's all-reduce bandwidth, measured with NVIDIA's nccl-tests. It builds against the CUDA toolkit and NCCL, both from NVIDIA's repository, which carries an NCCL build for each CUDA release. These steps use CUDA 12.8 on an 8-GPU VM:
-
Install the CUDA 12.8 toolkit as the driver and CUDA guide shows, which also adds NVIDIA's repository.
-
Install NCCL built for CUDA 12.8, with its headers, and the compilers:
sudo apt-get install -y build-essential libnccl2=2.26.2-1+cuda12.8 libnccl-dev=2.26.2-1+cuda12.8 -
Build nccl-tests and run an all-reduce across all eight GPUs:
git clone https://github.com/NVIDIA/nccl-tests.git cd nccl-tests make CUDA_HOME=/usr/local/cuda-12.8 ./build/all_reduce_perf -b 8 -e 4G -f 2 -g 8
Version 2.26.2 was the newest NCCL built for CUDA 12.8 in NVIDIA's Ubuntu 22.04 repository on 2026-09-28. Without a version, apt-get picks the newest NCCL, which is built for a newer CUDA release.
Read the busbw column at the largest sizes. nccl-tests defines bus bandwidth so that it reflects the speed of the hardware bottleneck, whether that is NVLink, PCIe or the link between CPU sockets, which makes it the number to compare between VMs. On a 2-GPU VM, use -g 2.
When one 8-GPU server is not enough#
Past one server, the NVLink domain ends at the edge of the HGX board or the NVL72 rack, and the rest is networking. On-demand QuantaCloud VMs stop at 8 GPUs in one machine. For more, we build to order: HGX B200 and B300 servers with NVSwitch and 14.4 TB/s of NVLink per board, GB200 or GB300 NVL72 racks with 72 GPUs in one NVLink domain on request (GB200), and InfiniBand or RoCE fabrics between servers. The B200, B300 and GPU cluster pages cover the options, and dedicated GPU servers covers single machines. The configuration, lead time and terms come back in writing when you send a capacity brief.
SXM, PCIe and NVLink questions#
What does SXM mean?
SXM is NVIDIA's module form factor for data-center GPUs. Instead of plugging into a PCIe slot, the GPU mounts flat on an HGX baseboard, which connects it to the other GPUs on the board over NVLink and allows a higher power limit than a card. The A100's module is SXM4 and the H100's is SXM5, and the B200 and B300 ship on HGX boards in the same way.
Is NVLink faster than PCIe?
Yes, by a wide margin on data-center GPUs. The H100's 900 GB/s of NVLink is 7 times its PCIe Gen5 slot by NVIDIA's count, and the A100's 600 GB/s is about 9 times its PCIe Gen4 slot. The exception is the RTX A6000's bridge, at 112.5 GB/s, which is less than twice a PCIe Gen4 slot.
Can I add NVLink to an L40S, L40, RTX 6000 Ada or RTX PRO 6000?
No. NVIDIA lists NVLink as not supported on the L40S, L40, RTX 6000 Ada and RTX PRO 6000 Server Edition, and the Workstation Edition's specifications list no NVLink either, so GPUs of those models in one VM always talk over PCIe. For one model across them, use pipeline parallelism, and for many small models, run one per GPU.
Does NVLink turn two GPUs into one big GPU?
No. Each GPU keeps its own memory, and your framework decides how to split the work across them. vLLM, for example, needs --tensor-parallel-size or --pipeline-parallel-size to spread one model over several GPUs. NVLink only makes the traffic between them faster.
Should I rent the A100 SXM4 or the A100 PCIe?
For one GPU, whichever is cheaper when you check: the published TFLOPS and the 80 GB are the same. For eight GPUs training one model together, the 8x SXM4 VM, after nvidia-smi topo -m confirms NV entries between the GPUs. The PCIe VMs suit inference and independent jobs, where each GPU mostly works on its own.
My rule is short: for one GPU, ignore the link and compare memory and price. For one model on several GPUs, read nvidia-smi topo -m first, use tensor parallelism only across NV entries, and use pipeline parallelism across PCIe. When the job needs more than one 8-GPU server, the link that matters next is the network, which is where InfiniBand vs NVLink picks up.