Dedicated capacity · built to order

NVIDIA GB200 NVL72

13.4 TB. Built around your workload.

GB200 NVL72 specs from NVIDIA, about 120 kW and liquid cooling per rack, dated price reports, and how QuantaCloud scopes rack-scale builds.

Dedicated hardwareConfiguration scoped with you
A plan for your workloadBuilt to order
Plan dedicated capacity
Configuration and commercial terms agreed in writing.
NVIDIA GB200 NVL72GRACE BLACKWELL
VRAMVRAMVRAMVRAMVRAMVRAMNVL72
GPU memory per rack.13.4 TB HBM3E
GPU memory per rack
13.4 TB HBM3E
Memory bandwidth
576 TB/s
Architecture
Grace Blackwell
Deployment
Built to order

The GB200 NVL72 is a rack, not a server: 72 Blackwell GPUs and 36 Grace CPUs in one liquid-cooled cabinet that NVIDIA puts at about 120 kW at full load. Its building block is the GB200 superchip, one Grace CPU joined to two Blackwell GPUs, and each GB200 GPU has 186 GB of HBM3E. QuantaCloud builds GB200 NVL72 capacity on request, as reserved capacity. We discuss rack-scale builds before we quote them, because power and cooling have to be settled before any hardware is ordered.

Rack-scale builds are discussed first#

A GB200 build starts with a conversation, not a price list. Send a brief that says what you will run (training or inference, and the model sizes), how many racks, how they connect to each other and to storage, where they should run and when you need them.

The configuration, lead time and terms come back in writing before you commit.

Talk to us about GB200

GB200 NVL72 specs#

The GB200 NVL72 puts 13.4 TB of HBM3E into one 72-GPU NVLink domain. These are NVIDIA's figures from the GB200 NVL72 page and the Blackwell datasheet (October 2025), read on 2026-09-28. FP8 and INT8 are sparse figures, and dense is half.

SpecGB200 NVL72 rackGB200 superchipPer Blackwell GPU
Configuration36 Grace CPUs, 72 Blackwell GPUs1 Grace CPU, 2 Blackwell GPUs
GPU memory13.4 TB HBM3E372 GB HBM3E186 GB HBM3E
GPU memory bandwidth576 TB/s16 TB/s8 TB/s
CPU memory17 TB LPDDR5X at 14 TB/sUp to 480 GB LPDDR5X at up to 512 GB/s
CPU cores2,592 Arm Neoverse V272 Arm Neoverse V2
FP4 Tensor Core1,440 PFLOPS sparse, 720 dense40 PFLOPS sparse, 20 dense20 PFLOPS sparse, 10 dense
FP8/FP6 Tensor Core720 PFLOPS sparse20 PFLOPS sparse10 PFLOPS sparse
INT8 Tensor Core720 POPS sparse20 POPS sparse10 POPS sparse
FP642,880 TFLOPS80 TFLOPS40 TFLOPS
NVLink130 TB/s across the rack3.6 TB/s1.8 TB/s, 5th generation
PowerAbout 120 kW per rack at full loadConfigurable up to 1,200 W
Compute capability10.0 (sm_100)

Each GPU in a GB200 is the same Blackwell silicon as a B200, with more headroom: 186 GB at 8 TB/s instead of 180 GB at 7.7 TB/s in an HGX server, and a power limit of 1,200 W instead of 1,000 W. The superchip links it to the Grace CPU over NVLink-C2C at 900 GB/s, so the GPUs can reach the Grace memory as well as their own.

What is inside the rack#

Each NVL72 rack holds 18 compute trays and 9 NVLink switch trays, each one rack unit tall. In NVIDIA's DGX GB200, a compute tray carries two Grace CPUs and four Blackwell GPUs, four 400 Gb/s ConnectX-7 ports for the cluster network, two BlueField-3 DPUs for storage and management traffic, four 3.84 TB NVMe drives as a data cache and a 1.92 TB boot drive. Each switch tray holds two NVSwitch chips, and a passive copper cable backplane at the back of the rack ties all 72 GPUs into one NVLink domain. Two top-of-rack switches carry management traffic, and power shelves feed every tray through a bus bar.

Power and cooling come before hardware#

About 120 kW per rack is the number that decides where a GB200 NVL72 can go. NVIDIA's DGX rack guide puts rack power consumption at about 120 kW, and its Mission Control documentation gives the same figure for a rack at full load, switches included. The rack carries eight power shelves, each with six 5.5 kW supplies and up to 33 kW of output. That is about as much power as eight DGX B200 servers (our calculation: 120 / 14.3 = 8.4). Cooling is hybrid: liquid runs through manifolds and into cold plates on every CPU and GPU, while fans air-cool the network cards and drives. The room therefore needs a coolant supply to the rack's manifolds as well as the power, and that is the first thing we settle with you.

What an NVL72 is for#

An NVL72 earns its complexity when one job needs more than eight GPUs talking at NVLink speed. Inside the rack, every GPU reaches the other 71 at 1.8 TB/s. Between eight-GPU servers, a DGX B200 gives each GPU a single 400 Gb/s network port. That gap matters when tensor or expert parallelism spans more than eight GPUs. The GB200 also keeps 40 TFLOPS of FP64 per GPU, 2,880 per rack, so it can run double-precision HPC as well, which the GB300 largely gives up.

If your model and its KV cache fit in one eight-GPU server, which is 1.44 TB on an HGX B200, a B200 server is much easier to place: NVIDIA's DGX B200 is air-cooled and draws about 14.3 kW. We build those as dedicated GPU servers, several of them over InfiniBand make a GPU cluster, and the InfiniBand vs NVLink guide covers which traffic belongs on which link.

What a GB200 NVL72 costs#

The published estimates put a GB200 NVL72 rack at about $3 million. HSBC analysts estimated about $3 million per rack and $60,000 to $70,000 per GB200 superchip, figures a Barron's writer posted and Tom's Hardware reported on 2024-05-14. A Wolfe Research note covered by Investing.com on 2026-01-30 cited reported prices of about $3 million per rack. On 2026-03-24, Tom's Hardware quoted a source with knowledge of the matter at $2.8 million to $3.4 million depending on configuration, and the same source warned that many quotes online are manufacturer prices without a proper warranty.

NVIDIA has never confirmed a list price for its NVL72 racks, prices vary by configuration, and none of these figures covers the facility work around the rack. QuantaCloud quotes GB200 builds in writing, after the rack-scale conversation.

Software on Grace and Blackwell#

The GB200's GPUs are the same sm_100 as a B200's. They need CUDA 12.8 or newer, an R570 or newer driver and, on Linux, NVIDIA's open kernel modules, and vLLM states CUDA 12.8 as its minimum for GB200. What changes is the host. Grace uses Arm Neoverse V2 cores, so every container image and Python wheel you deploy needs an arm64 (aarch64) build. vLLM, for example, ships aarch64 wheels alongside its x86_64 ones. NVIDIA's DGX GB200 runs DGX OS, Ubuntu, Red Hat Enterprise Linux or Rocky, and comes with NVIDIA Mission Control for operating the rack.

GB200 questions#

What is the GB200?

It is a superchip: one Grace CPU and two Blackwell GPUs joined by NVLink-C2C at 900 GB/s. NVIDIA ships 36 of them in each GB200 NVL72 rack, which is where the 72 GPUs come from.

How much memory does a GB200 have?

Each Blackwell GPU in a GB200 has 186 GB of HBM3E, so one superchip carries 372 GB plus up to 480 GB of LPDDR5X on its Grace CPU. A full NVL72 rack has 13.4 TB of HBM3E and 17 TB of LPDDR5X.

How much does a GB200 NVL72 cost?

About $3 million per rack: HSBC analysts estimated that in 2024, a Wolfe Research note cited reported prices at that level in January 2026, and a Tom's Hardware source put it at $2.8 million to $3.4 million by configuration in March 2026. QuantaCloud quotes each build in writing.

Can I rent a GB200 by the hour?

No. GB200 capacity is reserved and built on request. The on-demand catalog on the GPUs page tops out at the H200 NVL with 141 GB, which is where I would prototype while a rack is being scoped. Stopping an on-demand instance deletes its disk, so copy results off first.

Should I choose GB200 or GB300?

Choose GB300 when inference memory and dense FP4 matter most: 279 GB per GPU instead of 186, 1.5x the dense FP4, 2x the attention throughput and 800 Gb/s of networking per GPU, with a GPU power limit of 1,400 W instead of 1,200 W. Choose GB200 when you need FP64 or INT8. The GB300 page has its full specs and power figures.

The rule I follow: consider a GB200 NVL72 only when one job needs more than eight GPUs in one NVLink domain, or FP64 at rack scale, and only once the site for each rack has liquid cooling and about 120 kW of power settled. Otherwise, build eight-GPU B200 servers and connect them over InfiniBand. Either way, start with a brief on the reserved capacity page that covers the models, rack count, network, storage, location and start date.

Talk to us about GB200

Keep building

Choose your next step.