Dedicated capacity · built to order

NVIDIA B200

180 GB. Built around your workload.

What a B200 costs by NVIDIA's own estimate, HGX B200 specs, what fits in 180 GB, and B200 servers and clusters QuantaCloud builds to order.

Dedicated hardwareConfiguration scoped with you
A plan for your workloadBuilt to order
Plan dedicated capacity
Configuration and commercial terms agreed in writing.
NVIDIA B200BLACKWELL
VRAMVRAMVRAMVRAMVRAMVRAMBLACKWELL
Room for the work ahead.180 GB HBM3E
GPU memory
180 GB HBM3E
Memory bandwidth
7.7 TB/s
Architecture
Blackwell
Deployment
Built to order

The B200 has no list price. The closest figure from NVIDIA is Jensen Huang's estimate in March 2024 that a Blackwell GPU would cost $30,000 to $40,000, and NVIDIA ships the B200 in eight-GPU HGX B200 systems, so the number you actually pay is a server or cluster price that moves with the configuration. QuantaCloud does not rent the B200 by the hour. We build B200 servers and clusters to order: you send a capacity brief, and the configuration, lead time and terms come back in writing before you commit.

B200 servers built to your spec#

Every B200 build is ordered for you. You send the GPU count, network, storage, region, start date and term. QuantaCloud replies in writing with a configuration, a lead time and commercial terms, and we order and build the hardware to that spec once you accept. The reserved capacity page lists what a useful brief includes.

Plan a B200 build

NVIDIA B200 specs#

Each B200 in an HGX B200 server has 180 GB of HBM3E, and the eight GPUs share 14.4 TB/s of NVLink. These are NVIDIA's figures, from the Blackwell datasheet (October 2025) and the HGX and DGX B200 pages as read on 2026-09-28.

SpecPer B200 GPU (HGX B200)HGX B200 system (8 GPUs)
ArchitectureBlackwell, compute capability 10.0 (sm_100)8x Blackwell SXM
GPU memory180 GB HBM3E1.4 TB (DGX B200: 1,440 GB)
Memory bandwidth7.7 TB/s62 TB/s
FP4 Tensor Core18 PFLOPS sparse, 9 dense144 PFLOPS sparse, 72 dense
FP8/FP6 Tensor Core9 PFLOPS sparse, 4.5 dense72 PFLOPS sparse, 36 dense
FP16/BF16 Tensor Core4.5 PFLOPS sparse, 2.25 dense36 PFLOPS sparse, 18 dense
INT8 Tensor Core9 POPS sparse72 POPS sparse
FP6437 TFLOPS296 TFLOPS
NVLink5th generation, 1.8 TB/s per GPU14.4 TB/s total through NVSwitch
Host linkPCIe Gen5, 128 GB/s
Max powerConfigurable up to 1,000 WDGX B200: about 14.3 kW max
Networking0.8 TB/s (DGX B200: one 400 Gb/s ConnectX-7 port per GPU)
MIGUp to 7 instances

Three numbers in that table need a note. NVIDIA's architecture comparison lists the Blackwell chip at 192 GB, but an HGX B200 exposes 180 GB per GPU and a GB200 exposes 186 GB, so plan on 180. The headline FP4, FP8 and INT8 figures assume sparsity and dense is half, so I plan on the dense column unless the model was pruned for sparse Tensor Cores. The datasheet gives 7.7 TB/s of bandwidth per GPU, while the DGX B200 page lists 64 TB/s for eight GPUs, or 8 TB/s each. I use the datasheet figure.

What the B200 is for, and what fits in 180 GB#

The B200 is built for FP8 training and low-precision inference of large models. FP8 runs at 4.5 dense PFLOPS per GPU and FP4 at 9. NVIDIA's NVFP4 format stores an FP8 scale for every block of 16 values plus one scale per tensor, and NVIDIA says that keeps accuracy close to FP8, often within about 1%, while using about 1.8x less memory than FP8 and up to 3.5x less than FP16. vLLM runs NVFP4 natively on Blackwell and falls back to weight-only 4-bit on GPUs without FP4 kernels. The B200 also keeps 37 TFLOPS of FP64, so it still handles double-precision simulation, which the B300 largely gives up.

The one thing I always check is whether the weights and the KV cache fit. The table uses vLLM's default --gpu-memory-utilization of 0.92, which gives the engine 164.7 GiB of each B200 (our calculation: a B200 reports 183,359 MiB, and 183,359 / 1,024 x 0.92 = 164.7), and it counts weights and KV cache only, before activations.

Model and precisionWeightsLeft on one B200What that holds (our calculation)
Llama-3.3-70B, BF16141.1 GB (131.4 GiB)33.3 GiBAbout 109,000 tokens of BF16 KV cache, short of one full 131K context
Llama-3.3-70B, FP872.7 GB (67.7 GiB)97.1 GiBAbout 318,000 tokens: two full 131K contexts
gpt-oss-120b, MXFP465.2 GB on disk (60.8 GiB)104.0 GiBAbout 23 full 131K contexts at 4.5 GiB each
Mistral Large 3 (675B), FP8675 GBDoes not fit one GPUMistral lists FP8 on a single node of B200s, which holds 1.44 TB
Kimi K3, MXFP4About 1.56 TB on diskDoes not fit one GPUMore than one HGX B200 holds: two servers, or one HGX B300

The Llama rows use a KV cache of 320 KiB per token at BF16 (our calculation from the model config: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). Setting --kv-cache-dtype fp8 halves every KV figure, which is enough for one full 131K context next to BF16 weights. The KV cache guide walks through the formula for other models.

Servers and clusters we build#

The smallest B200 build is one eight-GPU server, because NVIDIA's server option for the B200 is the eight-GPU HGX B200. Inside it, NVSwitch connects every GPU to every other at 1.8 TB/s. Between servers, traffic goes over InfiniBand or RoCE, and NVIDIA's DGX B200 gives each GPU its own 400 Gb/s ConnectX-7 port. The DGX B200 is also air-cooled, takes 10 rack units and draws about 14.3 kW at most, so an eight-GPU B200 server can run on air, where a GB200 NVL72 rack needs liquid cooling.

BuildGPUsInterconnectWhere it is described
One single-tenant HGX B200 server8NVLink and NVSwitch inside the serverDedicated GPU servers
B200 cluster16 and up, in steps of 8NVLink inside each server, InfiniBand or RoCE between themGPU clusters
Rack-scale Blackwell72 per NVL72 rackOne 72-GPU NVLink domainGB200 NVL72, discussed before it is quoted

The InfiniBand vs NVLink guide explains which traffic uses which link. If your procurement requires NVIDIA's own DGX B200 rather than another HGX B200 server, say so in the brief. Storage, region and term are part of the quote.

What a B200 costs#

The B200 has no public list price, so every figure here is an estimate. The closest one from NVIDIA is Jensen Huang's: on 2024-03-19 he told CNBC that a Blackwell GPU would cost $30,000 to $40,000, and he later told CNBC that the cost is not just the chip but also designing data centers and integrating them. CNBC noted in the same report that NVIDIA does not reveal list prices and that what a buyer pays depends on volume and on whether the GPUs come as a complete system. Eight GPUs at that range come to $240,000 to $320,000 before the chassis, CPUs, memory, network cards, storage and support (our calculation: 8 x $30,000 and 8 x $40,000). Server and cluster prices vary with every one of those choices, and since NVIDIA does not publish list prices, this page stops at the 2024 range.

QuantaCloud prices a B200 build as a written quote. What moves it is already in the brief: GPU count, fabric, storage, region, start date and term. For rack-scale Blackwell, the GB200 page has the dated rack estimates.

Software readiness for the B200#

The B200 needs CUDA 12.8 or newer and an R570 or newer driver, and on Linux it runs only with NVIDIA's open kernel modules. PyTorch gained Blackwell kernels in 2.7.0, in its CUDA 12.8 wheels only, and the CUDA 12.6 wheels still have none. Since PyTorch 2.11, a plain pip install torch pulls CUDA 13.0 wheels, and vLLM 0.30.0's default wheel and Docker image are CUDA 13.0 builds too. Both need an R580 or newer driver, and vLLM keeps -cu129 images for hosts on older drivers. On a delivered B200 server, two commands tell you where you stand:

Terminal
nvidia-smi --query-gpu=name,compute_cap,driver_version --format=csv
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"

The first should report compute capability 10.0, and the architecture list should include sm_100. The driver and CUDA version guide explains each field.

Start on-demand while the build is in progress#

You can start today on an on-demand GPU while a B200 build is on order. The H200 NVL has the most memory in the on-demand catalog at 141 GB, enough for Llama-3.3-70B in FP8 with one full 131K context (our calculation: 72.7 + 42.9 = 115.6 GB). The RTX PRO 6000 Blackwell also runs FP4, but it is compute capability 12.0 where the B200 is 10.0, so any kernel you compile yourself needs a build for each. For the B200 against the H200 and the H100, see H200 vs B200 and B200 vs H100.

GPUMemoryFromAvailable now
H200 NVL-Not listedNo
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes

Prices checked 6 Oct 2026, 01:40 UTC

On-demand instances are Ubuntu VMs in our US regions, and stopping one deletes its disk, so copy results off first. The deploy guide covers the console steps, and the H200 NVL page and the RTX PRO 6000 Blackwell page show live configurations.

B200 questions#

Can I rent a B200 by the hour on QuantaCloud?

No. The B200 is reserved capacity that we build to order, and it is not in the on-demand catalog. For hourly work today, the largest on-demand GPU is the H200 NVL with 141 GB, listed with the rest on the GPUs page.

How much memory does a B200 have?

An HGX B200 exposes 180 GB of HBM3E per GPU, or 1.44 TB across its eight GPUs. The GB200 superchip exposes 186 GB per GPU. The 192 GB in NVIDIA's architecture comparison is the chip's maximum capacity, not what a server gives you.

What is the difference between HGX B200 and DGX B200?

HGX B200 is NVIDIA's eight-GPU platform that certified server makers build systems on. DGX B200 is NVIDIA's own complete system on that platform: two Xeon Platinum 8570 CPUs with 112 cores in total, 2 TB of system memory (configurable to 4 TB), eight 400 Gb/s ConnectX-7 ports, 10 rack units and about 14.3 kW. Tell us in the brief if you need DGX specifically.

Should I build on B200 or B300?

Build on B300 when memory or long-context inference is the constraint: 270 GB per GPU against 180, and 14 dense FP4 PFLOPS against 9. Stay on B200 for FP8 or BF16 training that already fits, and for FP64 or INT8 work. The B300 vs B200 comparison has the numbers side by side.

How long does a B200 build take?

Lead time depends on GPU supply when the order is placed, and QuantaCloud confirms it in writing before you commit.

The rule I follow is simple: build on B200 when the model and its KV cache fit in 180 GB per GPU, or 1.44 TB per server, and the work is FP8 or BF16 training or FP8 serving. When memory, long context or FP4 throughput is the constraint, look at the B300 instead. Send the GPU count, fabric, storage, region, start date and term, and prototype on an on-demand H200 NVL while the hardware is on order.

Send a B200 capacity brief

Keep building

Choose your next step.