The GB200 NVL72 is a rack, not a server: 72 Blackwell GPUs and 36 Grace CPUs in one liquid-cooled cabinet that NVIDIA puts at about 120 kW at full load. Its building block is the GB200 superchip, one Grace CPU joined to two Blackwell GPUs, and each GB200 GPU has 186 GB of HBM3E. QuantaCloud builds GB200 NVL72 capacity on request, as reserved capacity. We discuss rack-scale builds before we quote them, because power and cooling have to be settled before any hardware is ordered.
Rack-scale builds are discussed first#
A GB200 build starts with a conversation, not a price list. Send a brief that says what you will run (training or inference, and the model sizes), how many racks, how they connect to each other and to storage, where they should run and when you need them.
The configuration, lead time and terms come back in writing before you commit.
Talk to us about GB200GB200 NVL72 specs#
The GB200 NVL72 puts 13.4 TB of HBM3E into one 72-GPU NVLink domain. These are NVIDIA's figures from the GB200 NVL72 page and the Blackwell datasheet (October 2025), read on 2026-09-28. FP8 and INT8 are sparse figures, and dense is half.
| Spec | GB200 NVL72 rack | GB200 superchip | Per Blackwell GPU |
|---|---|---|---|
| Configuration | 36 Grace CPUs, 72 Blackwell GPUs | 1 Grace CPU, 2 Blackwell GPUs | |
| GPU memory | 13.4 TB HBM3E | 372 GB HBM3E | 186 GB HBM3E |
| GPU memory bandwidth | 576 TB/s | 16 TB/s | 8 TB/s |
| CPU memory | 17 TB LPDDR5X at 14 TB/s | Up to 480 GB LPDDR5X at up to 512 GB/s | |
| CPU cores | 2,592 Arm Neoverse V2 | 72 Arm Neoverse V2 | |
| FP4 Tensor Core | 1,440 PFLOPS sparse, 720 dense | 40 PFLOPS sparse, 20 dense | 20 PFLOPS sparse, 10 dense |
| FP8/FP6 Tensor Core | 720 PFLOPS sparse | 20 PFLOPS sparse | 10 PFLOPS sparse |
| INT8 Tensor Core | 720 POPS sparse | 20 POPS sparse | 10 POPS sparse |
| FP64 | 2,880 TFLOPS | 80 TFLOPS | 40 TFLOPS |
| NVLink | 130 TB/s across the rack | 3.6 TB/s | 1.8 TB/s, 5th generation |
| Power | About 120 kW per rack at full load | Configurable up to 1,200 W | |
| Compute capability | 10.0 (sm_100) |
Each GPU in a GB200 is the same Blackwell silicon as a B200, with more headroom: 186 GB at 8 TB/s instead of 180 GB at 7.7 TB/s in an HGX server, and a power limit of 1,200 W instead of 1,000 W. The superchip links it to the Grace CPU over NVLink-C2C at 900 GB/s, so the GPUs can reach the Grace memory as well as their own.
What is inside the rack#
Each NVL72 rack holds 18 compute trays and 9 NVLink switch trays, each one rack unit tall. In NVIDIA's DGX GB200, a compute tray carries two Grace CPUs and four Blackwell GPUs, four 400 Gb/s ConnectX-7 ports for the cluster network, two BlueField-3 DPUs for storage and management traffic, four 3.84 TB NVMe drives as a data cache and a 1.92 TB boot drive. Each switch tray holds two NVSwitch chips, and a passive copper cable backplane at the back of the rack ties all 72 GPUs into one NVLink domain. Two top-of-rack switches carry management traffic, and power shelves feed every tray through a bus bar.
Power and cooling come before hardware#
About 120 kW per rack is the number that decides where a GB200 NVL72 can go. NVIDIA's DGX rack guide puts rack power consumption at about 120 kW, and its Mission Control documentation gives the same figure for a rack at full load, switches included. The rack carries eight power shelves, each with six 5.5 kW supplies and up to 33 kW of output. That is about as much power as eight DGX B200 servers (our calculation: 120 / 14.3 = 8.4). Cooling is hybrid: liquid runs through manifolds and into cold plates on every CPU and GPU, while fans air-cool the network cards and drives. The room therefore needs a coolant supply to the rack's manifolds as well as the power, and that is the first thing we settle with you.
What an NVL72 is for#
An NVL72 earns its complexity when one job needs more than eight GPUs talking at NVLink speed. Inside the rack, every GPU reaches the other 71 at 1.8 TB/s. Between eight-GPU servers, a DGX B200 gives each GPU a single 400 Gb/s network port. That gap matters when tensor or expert parallelism spans more than eight GPUs. The GB200 also keeps 40 TFLOPS of FP64 per GPU, 2,880 per rack, so it can run double-precision HPC as well, which the GB300 largely gives up.
If your model and its KV cache fit in one eight-GPU server, which is 1.44 TB on an HGX B200, a B200 server is much easier to place: NVIDIA's DGX B200 is air-cooled and draws about 14.3 kW. We build those as dedicated GPU servers, several of them over InfiniBand make a GPU cluster, and the InfiniBand vs NVLink guide covers which traffic belongs on which link.
What a GB200 NVL72 costs#
The published estimates put a GB200 NVL72 rack at about $3 million. HSBC analysts estimated about $3 million per rack and $60,000 to $70,000 per GB200 superchip, figures a Barron's writer posted and Tom's Hardware reported on 2024-05-14. A Wolfe Research note covered by Investing.com on 2026-01-30 cited reported prices of about $3 million per rack. On 2026-03-24, Tom's Hardware quoted a source with knowledge of the matter at $2.8 million to $3.4 million depending on configuration, and the same source warned that many quotes online are manufacturer prices without a proper warranty.
NVIDIA has never confirmed a list price for its NVL72 racks, prices vary by configuration, and none of these figures covers the facility work around the rack. QuantaCloud quotes GB200 builds in writing, after the rack-scale conversation.
Software on Grace and Blackwell#
The GB200's GPUs are the same sm_100 as a B200's. They need CUDA 12.8 or newer, an R570 or newer driver and, on Linux, NVIDIA's open kernel modules, and vLLM states CUDA 12.8 as its minimum for GB200. What changes is the host. Grace uses Arm Neoverse V2 cores, so every container image and Python wheel you deploy needs an arm64 (aarch64) build. vLLM, for example, ships aarch64 wheels alongside its x86_64 ones. NVIDIA's DGX GB200 runs DGX OS, Ubuntu, Red Hat Enterprise Linux or Rocky, and comes with NVIDIA Mission Control for operating the rack.
GB200 questions#
What is the GB200?
It is a superchip: one Grace CPU and two Blackwell GPUs joined by NVLink-C2C at 900 GB/s. NVIDIA ships 36 of them in each GB200 NVL72 rack, which is where the 72 GPUs come from.
How much memory does a GB200 have?
Each Blackwell GPU in a GB200 has 186 GB of HBM3E, so one superchip carries 372 GB plus up to 480 GB of LPDDR5X on its Grace CPU. A full NVL72 rack has 13.4 TB of HBM3E and 17 TB of LPDDR5X.
How much does a GB200 NVL72 cost?
About $3 million per rack: HSBC analysts estimated that in 2024, a Wolfe Research note cited reported prices at that level in January 2026, and a Tom's Hardware source put it at $2.8 million to $3.4 million by configuration in March 2026. QuantaCloud quotes each build in writing.
Can I rent a GB200 by the hour?
No. GB200 capacity is reserved and built on request. The on-demand catalog on the GPUs page tops out at the H200 NVL with 141 GB, which is where I would prototype while a rack is being scoped. Stopping an on-demand instance deletes its disk, so copy results off first.
Should I choose GB200 or GB300?
Choose GB300 when inference memory and dense FP4 matter most: 279 GB per GPU instead of 186, 1.5x the dense FP4, 2x the attention throughput and 800 Gb/s of networking per GPU, with a GPU power limit of 1,400 W instead of 1,200 W. Choose GB200 when you need FP64 or INT8. The GB300 page has its full specs and power figures.
The rule I follow: consider a GB200 NVL72 only when one job needs more than eight GPUs in one NVLink domain, or FP64 at rack scale, and only once the site for each rack has liquid cooling and about 120 kW of power settled. Otherwise, build eight-GPU B200 servers and connect them over InfiniBand. Either way, start with a brief on the reserved capacity page that covers the models, rack count, network, storage, location and start date.
Talk to us about GB200