The B200 does the work of two to three H100s on paper. Per GPU it has 180 GB of HBM3E against 80 GB, 7.7 TB/s of memory bandwidth against 3.35 TB/s on the H100 SXM, 2.3 times the SXM's dense FP8 and BF16 Tensor Core math, FP4 that Hopper does not have, and fifth-generation NVLink at 1.8 TB/s. The H100 is the one you can rent today: QuantaCloud rents the H100 PCIe by the hour from the console price, and we build HGX H100 and HGX B200 servers to order as reserved capacity, with the configuration, lead time and terms in writing before you commit.
Launch H100 PCIe Plan a B200 buildB200 vs H100 specs#
Two H100s matter here: the H100 PCIe that QuantaCloud rents by the hour, and the H100 SXM on HGX boards that we build to order. The figures are NVIDIA's, from its H100 whitepaper and product brief, its HGX page and the Blackwell datasheet. Every Tensor Core row is the dense figure, half of NVIDIA's sparse headline number.
| Per GPU | H100 PCIe (on demand) | H100 SXM (built to order) | B200 (built to order) |
|---|---|---|---|
| Architecture | Hopper, compute capability 9.0 | Hopper, compute capability 9.0 | Blackwell, compute capability 10.0 |
| GPU memory | 80 GB HBM2e | 80 GB HBM3 | 180 GB HBM3E |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 7.7 TB/s |
| FP4 Tensor Core | No | No | 9 PFLOPS |
| FP8 Tensor Core | 1,513 TFLOPS | 1,979 TFLOPS | 4.5 PFLOPS |
| BF16 and FP16 Tensor Core | 756 TFLOPS | 989 TFLOPS | 2.25 PFLOPS |
| FP64 and FP64 Tensor Core | 25.6 and 51.2 TFLOPS | 33.5 and 66.9 TFLOPS | 37 and 37 TFLOPS |
| GPU to GPU | NVLink bridge to one adjacent card, 600 GB/s | Fourth-generation NVLink, 900 GB/s | Fifth-generation NVLink, 1.8 TB/s |
| Max power | 350 W | Up to 700 W | Up to 1,000 W |
| Multi-Instance GPU | Up to 7 x 10 GB | Up to 7 x 10 GB | Up to 7 instances |
| Server | QuantaCloud VMs with 1 or 2 GPUs | HGX H100, 4 or 8 GPUs | HGX B200, 8 GPUs |
| Memory in an eight-GPU server | Not offered | 640 GB | 1,440 GB |
Against the H100 SXM, the B200 has 2.25 times the memory, 2.3 times the bandwidth and 2.3 times the dense FP8 and BF16 (our calculation: 180 / 80, 7.7 / 3.35 and 4,500 / 1,979). Against the PCIe card you can rent by the hour, the gaps grow to 3.85 times the bandwidth and 3 times the dense math (our calculation: 7.7 / 2.0 and 4,500 / 1,513 = 2.97). FP64 is the exception: the H100 SXM's FP64 Tensor Cores reach 66.9 TFLOPS, and the B200 lists 37 for FP64 and FP64 Tensor Core alike.
NVIDIA's 3X and 15X, with their conditions#
The headline numbers are projections with narrow conditions. NVIDIA says the DGX B200 delivers 3X the training performance and 15X the inference performance of previous-generation systems, and its footnotes, which call both figures projected performance subject to change, compare it with the DGX H100. The 15X compares eight air-cooled DGX H100 systems with one air-cooled DGX B200, per GPU, at 50 ms between tokens, a 5-second first token, 32,768 input tokens and 1,028 output tokens. The 3X compares clusters at 32,768-GPU scale, 4,096 eight-GPU systems of each on a 400 Gb/s InfiniBand network.
I read the 15X as a ceiling for very large models served under a strict latency target, where a slower GPU has to cut its batch size to keep up. For everything else I plan on the spec ratios: about 2.3 times an H100 SXM on dense math and on bandwidth.
What fits in 80 GB and in 180 GB#
Memory is where one B200 replaces several H100s. At vLLM's default --gpu-memory-utilization of 0.92, the engine gets 73.3 GiB of an H100 and 164.7 GiB of a B200 (our calculation from the 81,559 and 183,359 MiB that nvidia-smi reports, x 0.92). The table counts weights and KV cache only, before activations.
| Model and precision | Weights | Left on one H100 | Left on one B200 |
|---|---|---|---|
| gpt-oss-120b, MXFP4 | 65.2 GB on disk (60.8 GiB) | 12.5 GiB: 2.8 full 131K contexts | 104.0 GiB: about 23 full 131K contexts |
| Qwen3-32B, FP8 | 34.3 GB (32.0 GiB) | 41.3 GiB: 5.2 full 32K contexts | 132.8 GiB: 16 full 32K contexts |
| Llama-3.3-70B, FP8 | 72.7 GB (67.7 GiB) | 5.6 GiB: about 18,000 tokens, for short contexts only | 97.1 GiB: about 318,000 tokens |
| Llama-3.3-70B, BF16 | 141.1 GB (131.4 GiB) | Needs two H100s | 33.3 GiB: about 109,000 tokens |
The KV figures come from each model's config: 320 KiB per token for Llama-3.3-70B and 256 KiB for Qwen3-32B at BF16, and 4.5 GiB (4.83 GB) per full 131K context for gpt-oss-120b. The KV cache guide has the formula for other models.
Per eight-GPU server, an HGX H100 holds 640 GB and an HGX B200 1,440 GB. Mistral lists its 675B Mistral Large 3 in FP8 for a single node of B200s, and its NVFP4 version for a single node of H100s, which Hopper runs weight-only. Meta says the FP8 weights of Llama 4 Maverick, 401.6 GB by parameter count, fit on a single DGX H100.
Inference: what the B200 changes#
For serving, the B200's gains stack up. One GPU holds models that need two H100s for long contexts, so a 70B model at FP8 runs without splitting it across GPUs. Single-user decoding follows memory bandwidth, and batched serving and long prompts follow dense FP8, and both are about 2.3 times an H100 SXM.
FP4 is the new part. The B200 runs NVIDIA's NVFP4 format natively at 9 dense PFLOPS, while vLLM runs NVFP4 checkpoints on an H100 as weight-only 4-bit through its Marlin kernels, which saves memory without the FP4 speed. NVIDIA says NVFP4 keeps accuracy close to FP8, often within about 1 percent, in about 1.8 times less memory than FP8. The NVFP4 guide explains the format.
Training: fewer GPUs for the same job#
For training, the B200 changes how many GPUs a job needs. Full fine-tuning with mixed-precision Adam keeps 16 bytes per parameter for weights, gradients and optimizer states. For an 8B model that is about 131 GB: more than one H100 holds, so the states must be sharded across at least two, while one B200 holds them with 57 GiB left for activations (our calculation: 8.19B x 16 bytes = 122.0 GiB, then 179.1 - 122.0, from the 183,359 MiB a B200 reports). For a 70B model it is 1.12 TB, which one HGX B200 server holds with 389 GiB to spare, where HGX H100 needs two servers with 1,274 GiB between them (our calculation: 70B x 16 bytes = 1,043.1 GiB, 8 x 179.06 - 1,043.1 and 16 x 79.65). The LoRA, QLoRA and full fine-tuning guide covers the cheaper methods.
Inside the server, NVLink doubles from 900 GB/s to 1.8 TB/s per GPU. Between servers, NVIDIA's DGX H100 and DGX B200 both give each GPU a 400 Gb/s ConnectX-7 port, so multi-node jobs gain from needing fewer GPUs, not from a faster fabric. The InfiniBand vs NVLink guide explains which traffic takes which path, and GPU clusters covers multi-node builds.
Power per GPU rises from 700 W on an H100 SXM to 1,000 W on a B200. NVIDIA lists a DGX H100 at 10.2 kW maximum in 8 rack units and a DGX B200 at about 14.3 kW in 10. Per kilowatt, the DGX B200 carries about 1.6 times the dense FP8 of a DGX H100 (our calculation: 8 x 4.5 = 36 PFLOPS over 14.3 kW, against 8 x 1.979 = 15.8 PFLOPS over 10.2 kW).
Software: what carries over from the H100#
Framework code that runs on an H100 runs on a B200 once the stack is new enough. The B200 needs CUDA 12.8 or newer, an R570 or newer driver and, on Linux, NVIDIA's open kernel modules, and PyTorch's CUDA 12.6 wheels have no Blackwell kernels at all. Your own kernels need an sm_100 build, and kernels built for sm_90a, the Hopper-only target, do not run. FlashAttention-3 requires an H100 or H800, while FlashAttention-4 targets both Hopper and Blackwell. The driver and CUDA version guide shows how to check a machine.
When the H100 is still the right GPU#
The H100 is still the right GPU when the job fits in 80 GB and uses FP8, or when it is FP64 matrix work. gpt-oss-120b fits one H100 PCIe with room for 2.8 full-length contexts, Qwen3-32B at FP8 serves five 32K contexts at once, and Unsloth puts a 70B QLoRA fine-tune at 41 GB. For jobs like these I would rent an H100 by the hour today rather than wait for any build.
The on-demand H100 is the PCIe card, one or two per VM, and QuantaCloud's catalog lists these offers as PCIe. I treat a 2x VM as two cards on PCIe unless nvidia-smi topo -m shows an NV entry between them, and on PCIe vLLM's advice is pipeline parallelism rather than tensor parallelism.
| GPU | Memory | From | Available now |
|---|---|---|---|
| H100 PCIe | - | Not listed | No |
| H200 NVL | - | Not listed | No |
| RTX PRO 6000 Blackwell | - | Not listed | No |
Prices checked 5 Oct 2026, 20:08 UTC
When the model needs more than 80 GB and you want to start before a B200 build arrives, the H200 NVL with 141 GB is the on-demand step up, and H200 vs B200 compares that pair. Stopping an on-demand instance terminates it and deletes its disk, so copy results off before you stop. The H100 page has the live configurations.
Availability and price#
Neither the B200 nor the H100 SXM is sold by the hour at QuantaCloud. We order and build HGX H100 and HGX B200 servers to your spec, from a single server to an InfiniBand cluster, and the lead time comes back in writing with the quote, before you commit.
NVIDIA publishes no list prices, so each build is quoted for its configuration and term. The H100 price guide works through buying an H100 against renting one, and the B200 page has the closest figure NVIDIA has given for a Blackwell GPU.
B200 vs H100 questions#
Is the B200 worth it over the H100?
On the numbers, yes, when a job needs more than 80 GB per GPU or several H100s working together, because one B200 matches two to three H100s on memory and dense math. When the job fits in 80 GB at FP8, an H100 by the hour gets it running today, without a build.
Can I rent a B200 or an H100 SXM by the hour?
No. Both are reserved capacity that we build to order. The H100 PCIe is on demand, and the other on-demand GPUs are on the GPUs page.
Does H100 code run on a B200?
Framework code does, with a CUDA 12.8 or newer build of PyTorch or vLLM. Custom CUDA kernels need recompiling for sm_100, and Hopper-only kernels such as FlashAttention-3 need their Blackwell counterparts.
How much power does a B200 server draw?
NVIDIA lists about 14.3 kW maximum for a DGX B200, against 10.2 kW for a DGX H100, and up to 1,000 W per B200 against 700 W per H100 SXM.
The rule I follow: stay on the H100 when the job fits in 80 GB at FP8 or runs FP64 matrix math, and rent the H100 PCIe by the hour for it. Build on B200 when a model needs two or more H100s just to hold it, when you want FP4, or when a training job would need twice the H100 servers for its model states. Send the model, precision, context length and GPU count in one brief, and prototype on an on-demand H100 PCIe or H200 NVL while it is quoted. More GPU pairs are on the comparisons page.
Plan a B200 build