GPU comparison

H200 NVL vs B200

H200 vs B200 on NVIDIA's figures: 141 vs 180 GB, FP8, FP4, NVLink 4 vs 5 and power, and when to start on an H200 NVL while a B200 build is quoted.

Faiz Ahmed10 min read
NVIDIA · Hopper

H200 NVL

Cloud GPU with full SSH access

141 GBHBM3e memory
Bandwidth
4.8 TB/s
FP8 compute
Native FP8
From / GPU-hrSee console
Explore H200 NVL
NVIDIA · Blackwell

B200

Dedicated hardware, built to order

180 GBHBM3E memory
Bandwidth
7.7 TB/s
FP8 compute
Native FP8
How you get itBuilt to order
Plan capacity

Prices checked

The B200 is the bigger and faster GPU on almost every line of NVIDIA's spec sheets, and the H200 is the one you can run today. Per GPU, the B200 has 180 GB of HBM3E at 7.7 TB/s against the H200's 141 GB at 4.8 TB/s, 2.3 to 2.7 times the dense FP8 and BF16 Tensor Core throughput, FP4 that Hopper does not have, and fifth-generation NVLink at 1.8 TB/s against 900 GB/s. QuantaCloud rents the H200 NVL by the hour from the console price, and we build HGX H200 and HGX B200 servers to order as reserved capacity, with the configuration, lead time and terms in writing before you commit.

Launch H200 NVL Plan a B200 or H200 build

H200 vs B200 specs#

The H200 comes in two versions and the B200 in one. These are NVIDIA's figures from its H200 and HGX product pages and the Blackwell datasheet, as read on 2026-09-28. NVIDIA quotes most Tensor Core figures with sparsity and dense is half, so every Tensor Core row below is the dense figure.

Per GPUH200 NVL (on demand)H200 SXM (built to order)B200 (built to order)
ArchitectureHopper, compute capability 9.0Hopper, compute capability 9.0Blackwell, compute capability 10.0
GPU memory141 GB HBM3e141 GB HBM3e180 GB HBM3E
Memory bandwidth4.8 TB/s4.8 TB/s7.7 TB/s
FP4 Tensor CoreNoNo9 PFLOPS
FP8 Tensor Core1,671 TFLOPS1,979 TFLOPS4.5 PFLOPS
BF16 and FP16 Tensor Core835 TFLOPS989 TFLOPS2.25 PFLOPS
FP64 and FP64 Tensor Core30 and 60 TFLOPS34 and 67 TFLOPS37 and 37 TFLOPS
GPU to GPU2- or 4-way NVLink bridge, 900 GB/s per GPUFourth-generation NVLink, 900 GB/s per GPUFifth-generation NVLink, 1.8 TB/s per GPU
Host linkPCIe Gen5, 128 GB/sPCIe Gen5, 128 GB/sPCIe Gen5, 128 GB/s
Max powerUp to 600 WUp to 700 WUp to 1,000 W
Multi-Instance GPUUp to 7 instancesUp to 7 instancesUp to 7 instances
NVIDIA server platformMGX H200 NVL, up to 8 GPUsHGX H200, 4 or 8 GPUsHGX B200, 8 GPUs
Memory in an eight-GPU server1,128 GB1,128 GB1,440 GB

Three rows decide most builds: memory, dense math and NVLink. The B200 has 28 percent more memory and 60 percent more bandwidth per GPU (our calculation: 180 / 141 = 1.28 and 7.7 / 4.8 = 1.6). Its dense FP8 and BF16 throughput is 2.3 times the H200 SXM's and 2.7 times the H200 NVL's (our calculation: 4,500 / 1,979 = 2.27 and 4,500 / 1,671 = 2.69), and its NVLink runs at twice the speed. FP64 is the row that goes the other way. With FP64 Tensor Cores the H200 SXM reaches 67 TFLOPS, while the B200 lists 37 for FP64 and FP64 Tensor Core alike, so I keep double-precision matrix work on Hopper.

What fits in 141 GB and in 180 GB#

Memory decides what runs on one GPU, so check it before compute. At vLLM's default --gpu-memory-utilization of 0.92, the engine gets 129.2 GiB of an H200 and 164.7 GiB of a B200 (our calculation from the 143,771 and 183,359 MiB that nvidia-smi reports, x 0.92). The table counts weights and KV cache only, before activations.

Model and precisionWeightsLeft on one H200Left on one B200
Llama-3.3-70B, BF16141.1 GB (131.4 GiB)Does not fit33.3 GiB: about 109,000 tokens of BF16 KV cache
Llama-3.3-70B, FP872.7 GB (67.7 GiB)61.5 GiB: about 200,000 tokens, one full 131K context97.1 GiB: about 318,000 tokens, two full 131K contexts
gpt-oss-120b, MXFP465.2 GB on disk (60.8 GiB)68.4 GiB: about 15 full 131K contexts104.0 GiB: about 23 full 131K contexts

The Llama rows use 320 KiB of KV cache per token at BF16 (our calculation from the model config: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). gpt-oss-120b needs only 4.5 GiB (4.83 GB) per full 131K context, because half of its layers use a 128-token sliding window. Setting --kv-cache-dtype fp8 halves every KV figure, and the KV cache guide has the formula for other models.

Per eight-GPU server, an HGX H200 holds 1,128 GB and an HGX B200 1,440 GB. Mistral lists its 675B Mistral Large 3 in FP8 for a single node of either, so that model does not decide between them. Kimi K3, at about 1.56 TB on disk, needs more than one server of either generation.

Inference: bandwidth, FP8 and FP4#

For inference, what the B200 changes depends on the batch. Generating tokens for one user is bound by memory bandwidth, so the ratio to plan on there is 7.7 against 4.8 TB/s, or 1.6 times. Prefill and heavily batched serving are bound by Tensor Core math, where the B200 has 2.3 to 2.7 times the dense FP8.

FP4 is the new part. The B200 runs NVIDIA's NVFP4 format natively at 9 dense PFLOPS, and NVIDIA says NVFP4 keeps accuracy close to FP8, often within about 1 percent, in about 1.8 times less memory than FP8. The H200 has no FP4 Tensor Cores. vLLM still loads NVFP4 checkpoints on Hopper, but it runs them weight-only through its Marlin kernels, which saves memory without the FP4 speed. The NVFP4 guide explains the format, and FP8 vs FP16 vs BF16 covers the precisions both GPUs share.

For training, the B200's gain is dense math and the link inside the server. Dense BF16 goes from 989 TFLOPS on an H200 SXM to 2.25 PFLOPS, and dense FP8 from 1,979 TFLOPS to 4.5 PFLOPS, so a job bound by Tensor Core math can run up to 2.3 times faster per GPU on paper. NVLink doubles from 900 GB/s to 1.8 TB/s per GPU, which matters most for tensor parallelism and for the gradient traffic of data parallelism inside a server.

Between servers the two are level. NVIDIA's DGX H200 and DGX B200 both give each GPU a 400 Gb/s ConnectX-7 port, so a multi-node job gets no faster fabric from the B200 alone. The InfiniBand vs NVLink guide explains which traffic takes which path, and GPU clusters covers multi-node builds.

The extra memory changes what one GPU can train. Full fine-tuning of an 8B model with mixed-precision Adam needs about 131 GB for weights, gradients and optimizer states, which leaves 18.4 GiB of an H200 for activations and 57 GiB of a B200 (our calculation: 8.19B parameters x 16 bytes, the ZeRO paper's figure, is 122.0 GiB, then 140.4 - 122.0 and 179.1 - 122.0, from the memory each card reports). One H200 detail matters here: the NVL cards bridge at most four GPUs together, while HGX H200 and HGX B200 servers connect all eight through NVSwitch, so eight-way tensor parallelism is an HGX job on either generation.

Power rises with Blackwell. A B200 is allowed up to 1,000 W against 700 W for an H200 SXM and 600 W for an H200 NVL. NVIDIA lists a DGX H200 at 10.2 kW maximum in 8 rack units and a DGX B200 at about 14.3 kW in 10. Per kilowatt the DGX B200 still carries about 1.6 times the dense FP8 of a DGX H200 (our calculation: 8 x 4.5 = 36 PFLOPS over 14.3 kW, against 8 x 1.979 = 15.8 PFLOPS over 10.2 kW).

Software: sm_90 against sm_100#

The B200 needs a newer software stack than the H200. Hopper has been supported since CUDA 11.8, so any current stack runs on an H200. The B200 needs CUDA 12.8 or newer and an R570 or newer driver, and on Linux it runs only with NVIDIA's open kernel modules. PyTorch gained Blackwell kernels in 2.7.0, in its CUDA 12.8 wheels, and its CUDA 12.6 wheels still have none.

Your own kernels need an sm_100 build, and code built for sm_90a, the Hopper-only target, does not run on a B200. FlashAttention-3 is the common example: its README lists the H100 and H800 as requirements, while FlashAttention-4 targets both Hopper and Blackwell. The driver and CUDA version guide shows how to confirm what a server runs.

Start on an H200 NVL while the B200 build is quoted#

The on-demand H200 NVL is where I would start while a B200 build is quoted, as long as the model fits in 141 GB per GPU. Llama-3.3-70B at FP8 with one full 131K context needs 115.6 GB (our calculation: 72.7 + 42.9), so it runs on one H200 NVL today. A current PyTorch or vLLM build with CUDA 12.8 or newer runs on both GPUs, so your code, FP8 checkpoints, data pipeline and evaluation move to the B200 build unchanged. Two things do not carry over: FP4 speed, which only Blackwell hardware shows, and kernels you compile yourself, which need an sm_100 build.

On 2026-09-27 the catalog listed 1x and 2x H200 NVL VMs in us-east-1 (Virginia), and a 2x VM gives 282 GB for the jobs that miss on one card. The H200 NVL supports an NVLink bridge, but I check the link on the VM with nvidia-smi topo -m before I plan tensor parallelism around it.

GPUMemoryFromAvailable now
H200 NVL-Not listedNo
RTX PRO 6000 Blackwell-Not listedNo
H100 PCIe80 GB$2.59/GPU-hrYes

Prices checked 5 Oct 2026, 21:18 UTC

To try NVFP4 checkpoints before the B200 arrives, the on-demand RTX PRO 6000 Blackwell runs FP4, but it is compute capability 12.0, not 10.0, so kernels compiled for sm_100 do not run on it. On-demand instances are VMs in US regions, and stopping one terminates it and deletes its disk, so copy checkpoints off before you stop. The H200 page has the live configurations.

Availability and price#

Neither the H200 SXM nor the B200 is sold by the hour at QuantaCloud. We order and build HGX H200 and HGX B200 servers to your spec, from a single server to an InfiniBand cluster. Lead time depends on GPU supply when the order is placed, and it comes back in writing with the quote, before you commit.

NVIDIA publishes no list price for either GPU, so each build is quoted for its configuration and term. The H200 price guide works through buying against renting, and the B200 page has the closest figure NVIDIA has given, Jensen Huang's 2024 estimate for a Blackwell GPU. For the next step up, B300 vs B200 compares the two Blackwell builds, and B200 vs H100 the older Hopper baseline.

H200 vs B200 questions#

Is the B200 faster than the H200 for every workload?

No. FP64 matrix work is the exception: the H200 SXM lists 67 TFLOPS with FP64 Tensor Cores against 37 on the B200. For a single user generating tokens, the gain follows the 1.6 times bandwidth ratio more than the 2.3 to 2.7 times compute ratio.

Can I rent a B200 by the hour on QuantaCloud?

No. The B200 is reserved capacity that we build to order. The H200 NVL is on demand, and the other on-demand GPUs are listed on the GPUs page.

The card supports a 2- or 4-way NVLink bridge at 900 GB/s per GPU, and QuantaCloud's catalog flags the H200 NVL offers as NVLink. I treat a 2x VM as PCIe-only until nvidia-smi topo -m shows an NV entry between the two GPUs, and on PCIe vLLM's advice is pipeline parallelism rather than tensor parallelism.

Can the H200 run FP4 models?

It can load them, not run them at FP4. vLLM runs NVFP4 checkpoints on Hopper as weight-only 4-bit through its Marlin kernels: the weights take less memory, and the math runs at higher precision.


The rule I follow: build on B200 when the model, its KV cache or its training state needs more than 141 GB per GPU, when you want FP4, or when an eight-GPU job is bound by Tensor Core math. Keep FP64 matrix work on Hopper, and start on an on-demand H200 NVL whenever the model fits in 141 GB, while the B200 brief is quoted. If you are not sure, send one brief with the model, precision, context length and GPU count, and ask for both an HGX H200 and an HGX B200 configuration. More GPU pairs are compared on the comparisons page.

Plan a B200 or H200 build

Keep building

Choose your next step.