Dedicated capacity · built to order

NVIDIA GB300 NVL72

20 TB. Built around your workload.

GB300 NVL72 specs from NVIDIA: 72 Blackwell Ultra GPUs, 20 TB of HBM3E, up to 142 kW per rack, plus dated price reports and reserved rack builds.

Dedicated hardwareConfiguration scoped with you
A plan for your workloadBuilt to order
Plan dedicated capacity
Configuration and commercial terms agreed in writing.
NVIDIA GB300 NVL72GRACE BLACKWELL ULTRA
VRAMVRAMVRAMVRAMVRAMVRAMNVL72
GPU memory per rack.20 TB HBM3E
GPU memory per rack
20 TB HBM3E
Memory bandwidth
Up to 576 TB/s
Architecture
Grace Blackwell Ultra
Deployment
Built to order

The GB300 NVL72 is a liquid-cooled rack of 72 Blackwell Ultra GPUs and 36 Grace CPUs, and NVIDIA's NVL72 reference architecture says a full rack can require up to 142 kW. It holds 20 TB of HBM3E in one NVLink domain, 279 GB per GPU. QuantaCloud builds GB300 NVL72 capacity on request, as reserved capacity, and we discuss rack-scale builds before we quote them. At this scale the facility decides what is possible before the hardware does.

Rack-scale builds are discussed first#

Every GB300 conversation starts with where the rack will run. Your brief should cover the workload (serving or training, and the models), the number of racks, how they connect to each other and to storage, the location and the start date.

The configuration, lead time and terms come back in writing before you commit.

Talk to us about GB300

GB300 NVL72 specs#

The GB300 NVL72 has 37 TB of fast memory: 20 TB of HBM3E on the GPUs and 17 TB of LPDDR5X on the Grace CPUs. These are NVIDIA's figures from the GB300 NVL72 page, the Blackwell Ultra datasheet (October 2025) and NVIDIA's NVL72 reference architecture, read on 2026-09-28. FP8 and INT8 are quoted with sparsity, and dense is half.

SpecGB300 NVL72 rackPer Blackwell Ultra GPU
Configuration72 Blackwell Ultra GPUs, 36 Grace CPUs
GPU memory20 TB HBM3E279 GB HBM3E
GPU memory bandwidthUp to 576 TB/s8 TB/s
CPU memory17 TB LPDDR5X at 14 TB/s
CPU cores2,592 Arm Neoverse V2
FP4 Tensor Core1,440 PFLOPS sparse, 1,080 dense20 PFLOPS sparse, 15 dense
FP8/FP6 Tensor Core720 PFLOPS sparse10 PFLOPS sparse
INT8 Tensor Core24 POPS sparse330 TOPS sparse
FP64100 TFLOPS1.3 TFLOPS
NVLink130 TB/s across the rack1.8 TB/s over 18 links, 5th generation
Host linkPCIe Gen6, 256 GB/s
NetworkingOne 800 Gb/s ConnectX-8 port per GPU
PowerUp to 142 kW per rackConfigurable up to 1,400 W
Compute capability10.3 (sm_103)

Two rows decide most GB300 projects. The 279 GB per GPU sets how much model and KV cache each GPU holds before you have to shard wider, and the up to 142 kW per rack sets where the rack can go at all.

Inside the rack, and the power it needs#

A GB300 NVL72 rack holds 18 compute trays and 9 NVLink switch trays. Each compute tray in NVIDIA's DGX GB300 carries two Grace CPUs and four Blackwell Ultra GPUs, four 800 Gb/s ConnectX-8 ports and a BlueField-3 DPU for storage and management traffic. Each switch tray holds two NVSwitch chips, and every GPU runs 18 NVLink links, one to each of those 18 switches, over the copper backplane.

Up to 142 kW per rack is the figure a facility has to plan for. NVIDIA's NVL72 reference architecture says a full GB300 rack can require that much and feeds it from eight power shelves of 33 kW, each with six 5.5 kW supplies. NVIDIA's DGX rack guide, written for both GB200 and GB300 racks, gives consumption of about 120 kW, so 142 kW is the ceiling to provision for. That is about ten DGX B300 servers' worth of power in one rack (our calculation: 142 / 14 = 10.1). The rack also smooths its own draw: energy storage in the power shelves charges when GPU demand drops and discharges when it spikes, and NVIDIA reports a 30% lower peak grid demand when training the Megatron LLM. NVIDIA calls the rack fully liquid-cooled, with coolant running through the manifolds to cold plates on every CPU and GPU, although its DGX rack guide still lists fans for the network cards and drives. The room needs a coolant supply to those manifolds before anything else is decided.

GB300 against GB200#

The GB300 NVL72 is the GB200 NVL72 rebuilt around Blackwell Ultra, and NVIDIA's DGX GB300 datasheet puts the gain at 1.5x the dense FP4 and 2x the attention performance of DGX GB200.

Per rack unless notedGB200 NVL72GB300 NVL72
Memory per GPU186 GB HBM3E279 GB HBM3E
HBM3E per rack13.4 TB20 TB
FP4, dense720 PFLOPS1,080 PFLOPS
FP8, sparse720 PFLOPS720 PFLOPS
INT8, sparse720 POPS24 POPS
FP642,880 TFLOPS100 TFLOPS
Network port per GPU (DGX)400 Gb/s ConnectX-7800 Gb/s ConnectX-8
Max power per GPU1,200 W1,400 W
Compute capability10.010.3

FP8 stays at 720 sparse PFLOPS per rack, so for FP8 training the GB300's gains are memory and networking, not Tensor Core math. FP64 and INT8 go the other way: if any part of the work is double-precision simulation or INT8 inference, the GB200 is the better rack. Rack power does not compare cleanly: NVIDIA's DGX rack guide gives about 120 kW for either rack, while its NVL72 reference architecture plans a GB300 rack for up to 142 kW.

What a GB300 NVL72 is for#

NVIDIA built the GB300 NVL72 for reasoning inference at scale, and the memory numbers show why. Even the largest open models use a fraction of the rack. Kimi K3 ships about 1.56 TB of weights, and vLLM's recipe asks for at least eight GB300 GPUs. DeepSeek's own example runs DeepSeek-V4-Pro, about 893 GB on disk, on a single 4x GB300 node. Qwen3.8-2.4T at FP8 is 2,446 GB of weights, about an eighth of the rack's 20 TB (our calculation: 2,446 / 20,000 = 0.12). The rack pays off when you serve many replicas or spread experts wide with all 72 GPUs on NVLink, or when you train across 72 GPUs without leaving it.

If the model and its KV cache fit in one eight-GPU server, 2.16 TB on an HGX B300, a B300 server is far easier to place: NVIDIA's DGX B300 is air-cooled and draws about 14 kW. We build those as dedicated GPU servers, or connect several into a GPU cluster over InfiniBand.

What a GB300 NVL72 costs#

The published GB300 NVL72 estimates disagree with each other. A Wolfe Research note covered by Investing.com on 2026-01-30 cited reported prices of about $4.3 million per GB300 NVL72, against about $3 million for a GB200 NVL72. On 2026-03-24, Tom's Hardware quoted a source with knowledge of the matter at $6 million to $6.5 million for an inference-optimized GB300 NVL72, and the same source warned that many quotes online are manufacturer prices without a proper warranty. I cannot reconcile the two from public information, so treat both as dated estimates.

NVIDIA has never confirmed a list price for its NVL72 racks, prices vary by configuration, and none of these figures covers the facility work around the rack. QuantaCloud quotes GB300 builds in writing, after the rack-scale conversation.

Software on Grace and Blackwell Ultra#

The GB300's GPUs are compute capability 10.3, like the B300's, and its host CPUs are Arm. The GPUs need CUDA 12.9 or newer and an R580 driver (NVIDIA first lists the GB300 in 580.82.07), and on Linux they run only with the open kernel modules. Kernels built for sm_100a do not run on them, so new builds target sm_103a or the sm_100f family. vLLM recommends CUDA 13 for GB300 and its CUDA 12.9 wheels skip 10.3, so on Grace you need an arm64 build that also targets CUDA 13. NVIDIA's DGX GB300 ships with Mission Control and DGX OS and supports Ubuntu. The driver and CUDA version guide shows how to confirm what a node is running.

GB300 questions#

What is the difference between GB300 and B300?

The GB300 pairs Blackwell Ultra GPUs with Grace CPUs on a superchip and ships in the NVL72 rack, with 279 GB per GPU and 72 GPUs in one NVLink domain. The B300 is the same GPU generation in eight-GPU HGX servers, such as the DGX B300 with Intel Xeon host CPUs, at 270 GB per GPU. The B300 page covers the server version.

How much power does a GB300 NVL72 need?

Plan for up to 142 kW for a full rack, according to NVIDIA's NVL72 reference architecture, delivered through eight 33 kW power shelves. NVIDIA's DGX rack guide gives consumption of about 120 kW, and the GPUs alone are allowed 100.8 kW between them (our calculation: 72 x 1,400 W).

How much does a GB300 NVL72 cost?

Published estimates run from about $4.3 million (reported prices cited in a Wolfe Research note, January 2026) to $6 million to $6.5 million for an inference-optimized rack (a Tom's Hardware source, March 2026), and prices vary with configuration. QuantaCloud quotes each build in writing.

Can I rent GB300 by the hour?

No. GB300 capacity is reserved and built on request. The on-demand catalog on the GPUs page tops out at the H200 NVL with 141 GB.

The rule I follow: choose a GB300 NVL72 when FP4 inference at long context is most of the work, when 72 GPUs on one NVLink domain will actually be used, and when the facility question has an answer, which means liquid cooling and up to 142 kW for each rack. If one eight-GPU server holds the model, build B300 servers and cluster them instead. Either way, the first step is a brief on the reserved capacity page with the models, rack count, network, storage, location and start date.

Talk to us about GB300

Keep building

Choose your next step.