QuantaCloud builds GPU clusters to order. You choose the GPU, the number of nodes, the fabric between them and the storage, and we order and build the hardware to that spec for your term. A GPU cluster, in the sense that matters for AI, is a set of GPU servers joined by a network fast enough for one training or inference job to use every GPU at once. In NVIDIA's own reference clusters for the H100, B200 and B300, each node is an 8-GPU server, with InfiniBand or RoCE Ethernet between the nodes.
Plan a clusterNeed a GPU today? Start on demand on a single VM with up to 8 GPUs while the cluster is built.
The node sets three limits#
The node you choose sets three limits: GPU memory per node, the size of the NVLink domain, and the network bandwidth each GPU gets to the rest of the cluster. These are NVIDIA's figures for its HGX platforms, its NVL72 racks and the DGX systems built on them.
| Node | GPU memory per node | NVLink inside the node | Compute network in NVIDIA's design | Page |
|---|---|---|---|---|
| HGX H100 | 8 x 80 GB = 640 GB | 900 GB/s per GPU, all 8 GPUs switched | 8 x 400 Gb/s, ConnectX-7 (DGX H100) | H100 |
| HGX H200 | 8 x 141 GB = 1,128 GB | 900 GB/s per GPU, all 8 GPUs switched | 8 x 400 Gb/s, ConnectX-7 (DGX H200) | H200 |
| HGX B200 | 8 x 180 GB = 1,440 GB | 1.8 TB/s per GPU, 14.4 TB/s in total | 8 x 400 Gb/s, ConnectX-7 (DGX B200) | B200 |
| HGX B300 | 8 x 270 GB = 2,160 GB | 1.8 TB/s per GPU, 14.4 TB/s in total | 8 x 800 Gb/s, ConnectX-8 (DGX B300) | B300 |
| GB200 NVL72 | 13.4 TB across 72 GPUs in one rack | 130 TB/s across all 72 GPUs | 72 x 400 Gb/s, ConnectX-7, one per GPU (DGX GB200) | GB200 |
| GB300 NVL72 | 20 TB across 72 GPUs in one rack | 130 TB/s across all 72 GPUs | 72 x 800 Gb/s, ConnectX-8, one per GPU | GB300 |
Clusters of PCIe servers are possible too, with cards such as the L40S or the RTX PRO 6000 Blackwell. Those cards have no NVLink, so every transfer between GPUs in a server crosses PCIe, and I would use them for work that splits cleanly: inference replicas, batch jobs, and data-parallel training of models that fit on one card.
How many nodes you need#
Memory sets the minimum, and time sets the rest. A full fine-tune with mixed-precision Adam keeps 16 to 18 bytes per parameter in weights, gradients and optimizer states before any activations. For a 70B model that is 1.12 to 1.26 TB (our calculation: 70 billion parameters x 16 to 18 bytes).
| Node | GPUs for 1.26 TB, 1,173.5 GiB (our calculation) | Whole nodes | Memory left for activations |
|---|---|---|---|
| HGX H100, 80 GB per GPU | 15 (1,173.5 / 79.65 = 14.7) | 2 | 100.9 GiB across 16 GPUs |
| HGX H200, 141 GB per GPU | 9 (1,173.5 / 140.4 = 8.4) | 2 | 1,072.9 GiB across 16 GPUs |
| HGX B200, 180 GB per GPU | 7 (1,173.5 / 179.1 = 6.6) | 1 | 259.0 GiB across 8 GPUs |
| HGX B300, 270 GB per GPU | 5 (1,173.5 / 268.6 = 4.4) | 1 | 975.2 GiB across 8 GPUs |
Two of those rows are tight. On H100s I would plan a third node, and on B200s I would test one node before counting on it: activations, communication buffers and memory fragmentation all come on top of the model states.
Beyond the minimum, extra nodes buy time. The Megatron-LM authors' rule is tensor parallelism up to the number of GPUs in one server, pipeline parallelism across servers, and data parallelism to scale further. Data parallelism exchanges gradients over the fabric on every step, and ZeRO stage 3, which shards weights, gradients and optimizer states across GPUs, moves 1.5 times the data of plain data parallelism, according to the ZeRO paper. That traffic is what the fabric has to carry.
Inference rarely needs a cluster. Llama 3.3 70B with FP8 weights and one full 128k-token context needs about 115.6 GB for weights and KV cache (our calculation: a 72.7 GB FP8 checkpoint plus 42.9 GB of KV cache), which fits on one H200. Multi-node inference is for models whose weights do not fit on one node, and vLLM's guidance for it is tensor parallelism across the GPUs in each node and pipeline parallelism across nodes.
InfiniBand or RoCE#
My default is InfiniBand, with RoCE for one specific case. Both fabrics use RDMA, and with GPUDirect RDMA the network card reads and writes GPU memory directly on each server. InfiniBand runs at 400 Gb/s per port (NDR, on NVIDIA Quantum-2 switches) or 800 Gb/s (XDR, on Quantum-X800). RoCE, RDMA over Converged Ethernet, runs the same kind of transfer over an Ethernet fabric, and NVIDIA's ConnectX-8 adapters speak both. NCCL, NVIDIA's library for moving data between GPUs, drives RoCE through the same InfiniBand verbs path, and NCCL_IB_GID_INDEX selects the address it uses in RoCE mode.
InfiniBand is the default because NVIDIA's DGX SuperPOD designs for H100, B200 and B300 use it for the compute fabric, and a cluster that matches a reference design has fewer unknowns. The case for RoCE is a cluster that has to sit inside an Ethernet network your team already runs: NVIDIA's Spectrum-X switches add adaptive routing and congestion control for it, and NVIDIA publishes a B300 design with a Spectrum-X compute fabric at 2 x 400 GbE per GPU.
Whichever you choose, ask for three things. The first is one compute port per GPU, as in the DGX B200 and B300. The second is a rail-aligned topology: in NVIDIA's H100 and B200 SuperPOD designs, the ports for matching GPUs in every node share a leaf switch, so traffic on a rail is one hop from the other 31 nodes in a 32-node unit. The third is storage and management on their own networks, off the compute fabric. InfiniBand vs NVLink explains why the fabric gives each GPU a ninth or less of the bandwidth NVLink gives it inside the node.
Storage#
Storage comes down to three numbers: dataset size, read speed and checkpoint size. For reference, NVIDIA's DGX B200 carries 8 x 3.84 TB of NVMe in each node for local data, about 30.7 TB (our calculation), and its B200 SuperPOD design sizes the storage network for more than 40 GB/s of I/O per node, on a fabric separate from GPU traffic. Tell us how much data you train on, how fast each node needs to read it, how large and how frequent your checkpoints are, and whether every node reads the same files. Those answers decide between local NVMe, a shared file system, or both.
What to send for a cluster#
A cluster brief needs seven answers.
| Field | What to tell us |
|---|---|
| Nodes and GPU | The node platform and the node count, or the model size and precision you need to train or serve |
| Fabric | InfiniBand or RoCE, the port speed per GPU (400 or 800 Gb/s), and whether you need GPUDirect RDMA |
| Storage | Dataset size, checkpoint size and frequency, and whether the nodes share files |
| Software | The OS, the driver and CUDA versions, and the scheduler you plan to run, such as Slurm or Kubernetes |
| Acceptance tests | What must pass before handover, such as an all_reduce_perf run from NVIDIA's nccl-tests across every node |
| Region, start and term | US by default, when you need the cluster, and for how long |
| Budget and security | A monthly range, and any compliance requirement. QuantaCloud is not SOC 2 or HIPAA certified |
What comes back is the same as for any reserved build: the configuration, the lead time and the commercial terms, in writing, before you commit. Reserved capacity walks through the steps.
Start on demand while we build#
You can start today on a single VM. QuantaCloud's on-demand instances are single VMs with up to 8 GPUs, and multi-node is a reserved build. On 2026-09-27 the catalog had 8-GPU VMs with A100 SXM4 80GB, L40S and RTX 6000 Ada GPUs. Use one to get the training loop, the data pipeline and checkpointing right on one node, so the cluster starts with code that already works.
Launch an 8x A100 SXM4 VMBefore you set a tensor-parallel size on a VM, run nvidia-smi topo -m to see whether its GPUs share NVLink or talk over PCIe. Copy checkpoints off before you stop: stopping an on-demand instance deletes its disk. The deploy guide covers launching from the console.
GPU cluster FAQ#
What is a GPU cluster?
A GPU cluster is a group of GPU servers connected by a fast network so that one job can use all of their GPUs. Inside each server, GPUs talk over NVLink or PCIe. Between servers, they talk over InfiniBand or RoCE Ethernet, for example through NVIDIA's NCCL library.
Can I rent a GPU cluster by the hour?
Not by the hour. A multi-node cluster is a reserved build for a term. What you can rent by the hour is a single on-demand VM with up to 8 GPUs, which the GPU catalog lists with live prices.
Can I rent an 8x H100 server?
Yes, as a reserved build. An HGX H100 node has 8 SXM GPUs with 80 GB each and NVLink at 900 GB/s per GPU, and it can run on its own as a dedicated server or as one node of a cluster. The H100 page lists the H100s you can launch by the hour today.
How much does a GPU cluster cost?
It is quoted per build, in writing, because the GPU, the node count, the fabric, the storage and the term all move the price. For the hardware side of the question, the H100 price guide compares buying H100s with renting them by the hour.
Can you build GB200 or GB300 NVL72 racks?
Yes, on request. Both are liquid-cooled racks with 72 GPUs in one NVLink domain at 130 TB/s, and we discuss the workload, power and cooling with you before quoting.
Plan a cluster#
My rule for clusters: if your model, its activations and your batch fit in one 8-GPU server, you need a dedicated server, not a cluster. Add nodes when memory forces you to, or when one server cannot finish the job in the time you have. Then send a brief with the node platform, the node count and the fabric, and the configuration, lead time and terms come back in writing.