NVLink and InfiniBand are not alternatives: NVLink connects the GPUs inside a server, and InfiniBand connects the servers to each other. A multi-node training cluster uses both. On an 8-GPU HGX board, NVLink gives each GPU 900 GB/s (Hopper) or 1.8 TB/s (Blackwell), switched so that any GPU reaches any other at full speed. Between servers, each InfiniBand port carries 400 or 800 Gb/s. Counted the same way, in bytes per second in each direction, that is a gap of 9 to 18 times per GPU in NVIDIA's own designs (our calculation, shown below), and it decides how you split a job across GPUs.
What each link connects#
Each link has one job: NVLink scales a server up, and InfiniBand or RoCE scales a cluster out.
| NVLink and NVSwitch | InfiniBand | RoCE Ethernet | |
|---|---|---|---|
| Connects | GPUs inside one server, or 72 GPUs in an NVL72 rack | Servers in a cluster | Servers in a cluster |
| Role | Scale-up | Scale-out | Scale-out |
| Current speed | 900 GB/s per GPU (Hopper), 1.8 TB/s (Blackwell) | 400 Gb/s per port (NDR), 800 Gb/s (XDR) | 400 or 800 Gb/s per port |
| Quoted as | Gigabytes per second, both directions added | Gigabits per second, each direction | Gigabits per second, each direction |
| NVIDIA hardware | NVLink ports on the GPU, NVSwitch chips on the board or in the rack | ConnectX adapters, Quantum switches | ConnectX adapters, Spectrum-X switches |
| Traffic it suits | Tensor parallelism, which communicates inside every layer | Data parallelism (gradients each step) and pipeline parallelism (activations between stages) | The same as InfiniBand |
RoCE, RDMA over Converged Ethernet, is the Ethernet route to the same result: the direct memory transfers InfiniBand does, carried over an Ethernet fabric. NVIDIA's ConnectX-8 adapters speak both, at up to 800 Gb/s.
NVLink and NVSwitch: inside the server#
NVLink is a direct GPU-to-GPU link, and NVSwitch turns it into a switched network. An H100 has 18 fourth-generation NVLink links for 900 GB/s in total, and each link carries 25 GB/s in each direction. NVIDIA puts that at 7 times the bandwidth of PCIe Gen5. Blackwell GPUs such as the B200 and B300 double it to 1.8 TB/s per GPU, still over 18 links.
On an HGX board, NVSwitch chips connect all 8 GPUs so that any pair talks at full NVLink speed: NVIDIA lists 14.4 TB/s of total NVLink bandwidth for the HGX B200 and the HGX B300. The GB200 NVL72 and GB300 NVL72 stretch one NVLink domain across a whole rack, with 72 GPUs and 130 TB/s in total.
PCIe cards get a smaller NVLink domain, or none. The H200 NVL joins 2 or 4 cards with NVLink bridges at 900 GB/s per GPU, the H100 PCIe and A100 PCIe pair up at 600 GB/s, and the RTX A6000 pairs at 112.5 GB/s. The L40S, L40, RTX 6000 Ada and RTX PRO 6000 Blackwell have no NVLink at all, so their GPUs talk over PCIe.
To see what a machine actually has, run this on it:
nvidia-smi topo -m
In the matrix, NV followed by a number means two GPUs share that many bonded NVLinks. PIX and PXB mean the path crosses one or several PCIe switches, PHB means it goes through the CPU's PCIe host bridge, NODE means it crosses host bridges within one NUMA node, and SYS means it crosses the link between NUMA nodes, usually between CPU sockets. QuantaCloud's on-demand instances are VMs with 1, 2, 4 or 8 GPUs, so check this on the VM before you choose a tensor-parallel size, rather than assuming NVLink is there. The SSH guide covers getting a shell on the VM.
InfiniBand and RoCE: between servers#
InfiniBand is the compute fabric in NVIDIA's own reference clusters. Its DGX SuperPOD designs for H100, B200 and B300 all use InfiniBand between nodes, and NVIDIA also publishes a B300 design on Spectrum-X Ethernet. Port speeds have doubled with each generation.
| Generation | Speed per port | NVIDIA switch | Where NVIDIA uses it |
|---|---|---|---|
| HDR | 200 Gb/s | Quantum QM8700, 40 ports | DGX A100 cluster ports |
| NDR | 400 Gb/s | Quantum-2, 64 ports, 51.2 Tb/s bidirectional | DGX H100, H200 and B200, with ConnectX-7 |
| XDR | 800 Gb/s | Quantum-X800, 144 ports | DGX B300 and GB300 NVL72, with ConnectX-8 |
NVIDIA's DGX systems give each GPU its own compute port. A DGX H100 or H200 has 8 single-port ConnectX-7 cards at up to 400 Gb/s for the compute fabric, with storage and management on separate cards, and a DGX B300 has 8 ConnectX-8 ports at up to 800 Gb/s.
GPUDirect RDMA is what makes those ports fast in practice. It gives the network card a direct path to GPU memory over PCIe, so data does not take a detour through CPU memory, and NVIDIA's documentation requires the GPU and the card to share an upstream PCIe root complex and says a path through PCIe switches alone performs best. That is why the placement of each NIC relative to its GPU matters, not just the port speed.
RoCE runs the same RDMA transfers over Ethernet. NCCL, NVIDIA's collective communication library, drives it through the same InfiniBand verbs transport, with NCCL_IB_GID_INDEX selecting the address it uses in RoCE mode. NVIDIA's Spectrum-X switches add adaptive routing and telemetry-based congestion control for AI traffic on Ethernet. InfiniBand switches can also take part in the math: NVIDIA SHARP runs reductions inside the switch, at version 3 on Quantum-2 and version 4 on Quantum-X800.
Why 900 GB/s and 400 Gb/s are further apart than they look#
The units hide the gap. NVLink is quoted in gigabytes per second with both directions added, and InfiniBand in gigabits per second for each direction. Two NVIDIA figures show it: each Hopper NVLink link carries 25 GB/s in each direction and 18 of them make the 900 GB/s total, while a Quantum-2 switch with 64 ports of 400 Gb/s is rated at 51.2 Tb/s bidirectional, which only adds up if 400 Gb/s is each way.
| Link | NVIDIA's figure | Each direction (our calculation) |
|---|---|---|
| NVLink, Hopper (H100, H200) | 900 GB/s per GPU | 900 / 2 = 450 GB/s |
| NVLink, Blackwell (B200, B300) | 1.8 TB/s per GPU | 1,800 / 2 = 900 GB/s |
| PCIe Gen5 x16 | 128 GB/s | 128 / 2 = 64 GB/s |
| InfiniBand NDR port | 400 Gb/s | 400 / 8 = 50 GB/s |
| InfiniBand XDR port | 800 Gb/s | 800 / 8 = 100 GB/s |
With one port per GPU, an H100 has 450 GB/s each way to the GPUs in its own server and 50 GB/s each way to everything outside it, a ninth as much. A B200 with a 400 Gb/s ConnectX-7 port, as in the DGX B200, gets an eighteenth, and a B300 with an 800 Gb/s ConnectX-8 port gets a ninth (our calculation: 450 / 50 = 9, 900 / 50 = 18 and 900 / 100 = 9). NVIDIA's HGX table shows the same ratios at node level: 14.4 TB/s of NVLink against 0.8 TB/s of networking on the HGX B200 (18 to 1) and 1.6 TB/s on the HGX B300 (9 to 1). There both figures add the two directions, since eight 400 Gb/s ports make 0.8 TB/s only when both directions are counted (our calculation: 8 x 400 x 2 = 6,400 Gb/s = 0.8 TB/s, and 8 x 800 x 2 = 12,800 Gb/s = 1.6 TB/s).
What this means for training#
Training puts its heaviest traffic on NVLink and the rest on the fabric. Tensor parallelism splits each layer across GPUs and runs all-reduce operations inside every layer, so it belongs inside one NVLink domain. The Megatron-LM authors turned this into a rule: tensor parallelism up to the number of GPUs in a server, pipeline parallelism across servers, and data parallelism to scale further. Pipeline stages pass activations point to point, which a fabric handles well. Data parallelism exchanges gradients on every step, and ZeRO stage 3 moves 1.5 times the data of plain data parallelism, according to the ZeRO paper, so a sharded job leans on the fabric harder than a replicated one.
A worked case shows where the line falls. A full fine-tune of a 70B model needs 1.12 to 1.26 TB for weights, gradients and Adam states before activations (our calculation at 16 to 18 bytes per parameter). That is at least 15 H100s, which means two HGX servers and a fabric between them, or a single 8-GPU B300 server with 2,160 GB, where the whole job stays on NVLink.
The one thing I always check on a training cluster is the tensor-parallel size: it should never exceed the NVLink domain, which is 8 GPUs on an HGX board and 72 on an NVL72 rack.
What this means for inference#
Most inference never touches InfiniBand. Llama 3.3 70B with FP8 weights and a full 128k-token context needs about 115.6 GB for weights and KV cache (our calculation: a 72.7 GB FP8 checkpoint plus 42.9 GB of KV cache), which fits on one H200 with its 141 GB. Inside one server, vLLM runs tensor parallelism over NVLink, and on GPUs without NVLink, such as the L40S, its docs recommend pipeline parallelism instead. The KV cache guide and how much VRAM you need cover the sizing side.
Multi-node inference is for models whose weights do not fit on one server. vLLM's guidance there is tensor parallelism across the GPUs in each node and pipeline parallelism across nodes, with NCCL_IB_HCA pointing NCCL at the InfiniBand adapters. Its docs also note that traffic between nodes is not encrypted, so keep the fabric private to your cluster. Mixture-of-experts models add expert parallelism, which spreads experts across GPUs, and when those GPUs sit in different servers, the fabric carries that traffic too.
QuantaCloud's on-demand VMs are single machines: 1, 2, 4 or 8 GPUs in one VM and no multi-node option, so only the link inside the VM matters there.
What to ask for in a cluster build#
The brief should name both halves: the NVLink domain and the fabric.
| Ask for | Why |
|---|---|
| The NVLink domain: an 8-GPU HGX board with NVSwitch, or an NVL72 rack | It caps your tensor-parallel size |
| One compute port per GPU, at 400 Gb/s (NDR) or 800 Gb/s (XDR) | It is how NVIDIA builds the DGX B200 and DGX B300 |
| A rail-aligned topology, and the oversubscription ratio if it is not non-blocking | NVIDIA's H100 and B200 SuperPOD designs keep matching GPUs one hop apart across 32 nodes |
| GPUDirect RDMA, with each NIC under the same PCIe switch as its GPU | It removes the copy through CPU memory |
| Storage and management on separate networks | NVIDIA's SuperPOD designs split compute, storage, in-band and out-of-band networks |
An all_reduce_perf run from NVIDIA's nccl-tests across every node, with the bus bandwidth it reports | It shows the fabric performs as built before your job depends on it |
QuantaCloud builds both halves as reserved capacity: HGX servers with NVSwitch, NVL72 racks on request, and InfiniBand or RoCE fabrics between servers. GPU clusters covers node platforms and sizing, and dedicated GPU servers covers single machines.
The rule I follow: size NVLink for the largest tensor-parallel group, and size the fabric for everything that crosses servers. If one 8-GPU server holds your model, its activations and your batch, you do not need InfiniBand at all. If it does not, plan the cluster around the fabric and send a cluster brief: the configuration, lead time and terms come back in writing.