The B200 has no list price. The closest figure from NVIDIA is Jensen Huang's estimate in March 2024 that a Blackwell GPU would cost $30,000 to $40,000, and NVIDIA ships the B200 in eight-GPU HGX B200 systems, so the number you actually pay is a server or cluster price that moves with the configuration. QuantaCloud does not rent the B200 by the hour. We build B200 servers and clusters to order: you send a capacity brief, and the configuration, lead time and terms come back in writing before you commit.
B200 servers built to your spec#
Every B200 build is ordered for you. You send the GPU count, network, storage, region, start date and term. QuantaCloud replies in writing with a configuration, a lead time and commercial terms, and we order and build the hardware to that spec once you accept. The reserved capacity page lists what a useful brief includes.
Plan a B200 buildNVIDIA B200 specs#
Each B200 in an HGX B200 server has 180 GB of HBM3E, and the eight GPUs share 14.4 TB/s of NVLink. These are NVIDIA's figures, from the Blackwell datasheet (October 2025) and the HGX and DGX B200 pages as read on 2026-09-28.
| Spec | Per B200 GPU (HGX B200) | HGX B200 system (8 GPUs) |
|---|---|---|
| Architecture | Blackwell, compute capability 10.0 (sm_100) | 8x Blackwell SXM |
| GPU memory | 180 GB HBM3E | 1.4 TB (DGX B200: 1,440 GB) |
| Memory bandwidth | 7.7 TB/s | 62 TB/s |
| FP4 Tensor Core | 18 PFLOPS sparse, 9 dense | 144 PFLOPS sparse, 72 dense |
| FP8/FP6 Tensor Core | 9 PFLOPS sparse, 4.5 dense | 72 PFLOPS sparse, 36 dense |
| FP16/BF16 Tensor Core | 4.5 PFLOPS sparse, 2.25 dense | 36 PFLOPS sparse, 18 dense |
| INT8 Tensor Core | 9 POPS sparse | 72 POPS sparse |
| FP64 | 37 TFLOPS | 296 TFLOPS |
| NVLink | 5th generation, 1.8 TB/s per GPU | 14.4 TB/s total through NVSwitch |
| Host link | PCIe Gen5, 128 GB/s | |
| Max power | Configurable up to 1,000 W | DGX B200: about 14.3 kW max |
| Networking | 0.8 TB/s (DGX B200: one 400 Gb/s ConnectX-7 port per GPU) | |
| MIG | Up to 7 instances |
Three numbers in that table need a note. NVIDIA's architecture comparison lists the Blackwell chip at 192 GB, but an HGX B200 exposes 180 GB per GPU and a GB200 exposes 186 GB, so plan on 180. The headline FP4, FP8 and INT8 figures assume sparsity and dense is half, so I plan on the dense column unless the model was pruned for sparse Tensor Cores. The datasheet gives 7.7 TB/s of bandwidth per GPU, while the DGX B200 page lists 64 TB/s for eight GPUs, or 8 TB/s each. I use the datasheet figure.
What the B200 is for, and what fits in 180 GB#
The B200 is built for FP8 training and low-precision inference of large models. FP8 runs at 4.5 dense PFLOPS per GPU and FP4 at 9. NVIDIA's NVFP4 format stores an FP8 scale for every block of 16 values plus one scale per tensor, and NVIDIA says that keeps accuracy close to FP8, often within about 1%, while using about 1.8x less memory than FP8 and up to 3.5x less than FP16. vLLM runs NVFP4 natively on Blackwell and falls back to weight-only 4-bit on GPUs without FP4 kernels. The B200 also keeps 37 TFLOPS of FP64, so it still handles double-precision simulation, which the B300 largely gives up.
The one thing I always check is whether the weights and the KV cache fit. The table uses vLLM's default --gpu-memory-utilization of 0.92, which gives the engine 164.7 GiB of each B200 (our calculation: a B200 reports 183,359 MiB, and 183,359 / 1,024 x 0.92 = 164.7), and it counts weights and KV cache only, before activations.
| Model and precision | Weights | Left on one B200 | What that holds (our calculation) |
|---|---|---|---|
| Llama-3.3-70B, BF16 | 141.1 GB (131.4 GiB) | 33.3 GiB | About 109,000 tokens of BF16 KV cache, short of one full 131K context |
| Llama-3.3-70B, FP8 | 72.7 GB (67.7 GiB) | 97.1 GiB | About 318,000 tokens: two full 131K contexts |
| gpt-oss-120b, MXFP4 | 65.2 GB on disk (60.8 GiB) | 104.0 GiB | About 23 full 131K contexts at 4.5 GiB each |
| Mistral Large 3 (675B), FP8 | 675 GB | Does not fit one GPU | Mistral lists FP8 on a single node of B200s, which holds 1.44 TB |
| Kimi K3, MXFP4 | About 1.56 TB on disk | Does not fit one GPU | More than one HGX B200 holds: two servers, or one HGX B300 |
The Llama rows use a KV cache of 320 KiB per token at BF16 (our calculation from the model config: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). Setting --kv-cache-dtype fp8 halves every KV figure, which is enough for one full 131K context next to BF16 weights. The KV cache guide walks through the formula for other models.
Servers and clusters we build#
The smallest B200 build is one eight-GPU server, because NVIDIA's server option for the B200 is the eight-GPU HGX B200. Inside it, NVSwitch connects every GPU to every other at 1.8 TB/s. Between servers, traffic goes over InfiniBand or RoCE, and NVIDIA's DGX B200 gives each GPU its own 400 Gb/s ConnectX-7 port. The DGX B200 is also air-cooled, takes 10 rack units and draws about 14.3 kW at most, so an eight-GPU B200 server can run on air, where a GB200 NVL72 rack needs liquid cooling.
| Build | GPUs | Interconnect | Where it is described |
|---|---|---|---|
| One single-tenant HGX B200 server | 8 | NVLink and NVSwitch inside the server | Dedicated GPU servers |
| B200 cluster | 16 and up, in steps of 8 | NVLink inside each server, InfiniBand or RoCE between them | GPU clusters |
| Rack-scale Blackwell | 72 per NVL72 rack | One 72-GPU NVLink domain | GB200 NVL72, discussed before it is quoted |
The InfiniBand vs NVLink guide explains which traffic uses which link. If your procurement requires NVIDIA's own DGX B200 rather than another HGX B200 server, say so in the brief. Storage, region and term are part of the quote.
What a B200 costs#
The B200 has no public list price, so every figure here is an estimate. The closest one from NVIDIA is Jensen Huang's: on 2024-03-19 he told CNBC that a Blackwell GPU would cost $30,000 to $40,000, and he later told CNBC that the cost is not just the chip but also designing data centers and integrating them. CNBC noted in the same report that NVIDIA does not reveal list prices and that what a buyer pays depends on volume and on whether the GPUs come as a complete system. Eight GPUs at that range come to $240,000 to $320,000 before the chassis, CPUs, memory, network cards, storage and support (our calculation: 8 x $30,000 and 8 x $40,000). Server and cluster prices vary with every one of those choices, and since NVIDIA does not publish list prices, this page stops at the 2024 range.
QuantaCloud prices a B200 build as a written quote. What moves it is already in the brief: GPU count, fabric, storage, region, start date and term. For rack-scale Blackwell, the GB200 page has the dated rack estimates.
Software readiness for the B200#
The B200 needs CUDA 12.8 or newer and an R570 or newer driver, and on Linux it runs only with NVIDIA's open kernel modules. PyTorch gained Blackwell kernels in 2.7.0, in its CUDA 12.8 wheels only, and the CUDA 12.6 wheels still have none. Since PyTorch 2.11, a plain pip install torch pulls CUDA 13.0 wheels, and vLLM 0.30.0's default wheel and Docker image are CUDA 13.0 builds too. Both need an R580 or newer driver, and vLLM keeps -cu129 images for hosts on older drivers. On a delivered B200 server, two commands tell you where you stand:
nvidia-smi --query-gpu=name,compute_cap,driver_version --format=csv
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"
The first should report compute capability 10.0, and the architecture list should include sm_100. The driver and CUDA version guide explains each field.
Start on-demand while the build is in progress#
You can start today on an on-demand GPU while a B200 build is on order. The H200 NVL has the most memory in the on-demand catalog at 141 GB, enough for Llama-3.3-70B in FP8 with one full 131K context (our calculation: 72.7 + 42.9 = 115.6 GB). The RTX PRO 6000 Blackwell also runs FP4, but it is compute capability 12.0 where the B200 is 10.0, so any kernel you compile yourself needs a build for each. For the B200 against the H200 and the H100, see H200 vs B200 and B200 vs H100.
| GPU | Memory | From | Available now |
|---|---|---|---|
| H200 NVL | - | Not listed | No |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
Prices checked 6 Oct 2026, 01:40 UTC
On-demand instances are Ubuntu VMs in our US regions, and stopping one deletes its disk, so copy results off first. The deploy guide covers the console steps, and the H200 NVL page and the RTX PRO 6000 Blackwell page show live configurations.
B200 questions#
Can I rent a B200 by the hour on QuantaCloud?
No. The B200 is reserved capacity that we build to order, and it is not in the on-demand catalog. For hourly work today, the largest on-demand GPU is the H200 NVL with 141 GB, listed with the rest on the GPUs page.
How much memory does a B200 have?
An HGX B200 exposes 180 GB of HBM3E per GPU, or 1.44 TB across its eight GPUs. The GB200 superchip exposes 186 GB per GPU. The 192 GB in NVIDIA's architecture comparison is the chip's maximum capacity, not what a server gives you.
What is the difference between HGX B200 and DGX B200?
HGX B200 is NVIDIA's eight-GPU platform that certified server makers build systems on. DGX B200 is NVIDIA's own complete system on that platform: two Xeon Platinum 8570 CPUs with 112 cores in total, 2 TB of system memory (configurable to 4 TB), eight 400 Gb/s ConnectX-7 ports, 10 rack units and about 14.3 kW. Tell us in the brief if you need DGX specifically.
Should I build on B200 or B300?
Build on B300 when memory or long-context inference is the constraint: 270 GB per GPU against 180, and 14 dense FP4 PFLOPS against 9. Stay on B200 for FP8 or BF16 training that already fits, and for FP64 or INT8 work. The B300 vs B200 comparison has the numbers side by side.
How long does a B200 build take?
Lead time depends on GPU supply when the order is placed, and QuantaCloud confirms it in writing before you commit.
The rule I follow is simple: build on B200 when the model and its KV cache fit in 180 GB per GPU, or 1.44 TB per server, and the work is FP8 or BF16 training or FP8 serving. When memory, long context or FP4 throughput is the constraint, look at the B300 instead. Send the GPU count, fabric, storage, region, start date and term, and prototype on an on-demand H200 NVL while the hardware is on order.
Send a B200 capacity brief