The B300 is a B200 with 50% more memory and about 50% more dense FP4, and it gives up most of its FP64 and INT8. Per GPU in an HGX server, that is 270 GB against 180 GB and 14 dense FP4 PFLOPS against 9, while FP8, BF16, memory bandwidth and NVLink stay where they were. For inference at long context, the B300 is the better build. For FP8 or BF16 training that already fits in 180 GB, the B200 does the same Tensor Core math on paper. QuantaCloud builds both to order as reserved capacity, and neither is sold by the hour.
B300 vs B200 specs#
Bandwidth, FP8, BF16 and NVLink are identical, and nearly everything else moves. These are NVIDIA's figures, from the Blackwell and Blackwell Ultra datasheets (both October 2025) and the HGX and DGX pages as read on 2026-09-28. Dense figures are half the sparse ones, except FP4, which NVIDIA lists separately.
| Per GPU in an HGX server | B200 | B300 | What changes |
|---|---|---|---|
| Architecture | Blackwell, compute capability 10.0 (sm_100) | Blackwell Ultra, compute capability 10.3 (sm_103) | New target for compiled kernels |
| GPU memory | 180 GB HBM3E | 270 GB HBM3E | 1.5x |
| Memory bandwidth | 7.7 TB/s | 7.7 TB/s | Same |
| FP4 Tensor Core | 18 PFLOPS sparse, 9 dense | 18 PFLOPS sparse, 14 dense | Dense about 1.5x, sparse the same |
| FP8/FP6 Tensor Core | 9 PFLOPS sparse, 4.5 dense | 9 PFLOPS sparse, 4.5 dense | Same |
| FP16/BF16 Tensor Core | 4.5 PFLOPS sparse, 2.25 dense | 4.5 PFLOPS sparse, 2.25 dense | Same |
| INT8 Tensor Core | 9 POPS sparse | 307 TOPS sparse | Cut |
| FP64 | 37 TFLOPS | 1.2 TFLOPS | Cut |
| Attention layers | Baseline | Up to 2x faster, per NVIDIA | SFU throughput for softmax doubled |
| NVLink | 5th generation, 1.8 TB/s | 5th generation, 1.8 TB/s | Same |
| Host link | PCIe Gen5, 128 GB/s | PCIe Gen6, 256 GB/s | 2x |
| Network port per GPU (DGX) | 400 Gb/s ConnectX-7 | 800 Gb/s ConnectX-8 | 2x |
| Max power | Up to 1,000 W | Up to 1,100 W | 100 W more |
| Eight-GPU server memory | 1.44 TB | 2.16 TB (NVIDIA: 2.1 TB) | 1.5x |
Memory: 180 GB against 270 GB#
Memory is the first difference to check, because it decides what fits on one GPU and in one server. At vLLM's default --gpu-memory-utilization of 0.92, the engine gets 164.7 GiB of a B200 and 247.1 GiB of a B300 (our calculation from the 183,359 and 275,040 MiB that nvidia-smi reports, x 0.92). The table counts weights and KV cache only, before activations.
| Model and precision | Weights | Left on one B200 | Left on one B300 |
|---|---|---|---|
| Llama-3.3-70B, BF16 | 141.1 GB (131.4 GiB) | 33.3 GiB: about 109,000 tokens of BF16 KV cache, short of one full 131K context | 115.7 GiB: about 379,000 tokens, nearly three full contexts |
| Llama-3.3-70B, FP8 | 72.7 GB (67.7 GiB) | 97.1 GiB: about 318,000 tokens | 179.4 GiB: about 588,000 tokens |
| Kimi K3, MXFP4 | About 1.56 TB on disk | Needs more than one HGX B200 (1.44 TB) | Fits one HGX B300 (2.16 TB) |
The Llama rows use 320 KiB of KV cache per token at BF16 (our calculation: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). DeepSeek-V4-Flash makes the same point from the other side: its FP4 and FP8 weights take about 167 GB on disk (155.5 GiB), which leaves only 9.2 GiB of the 164.7 GiB vLLM gets of a B200 at its default setting, while a B300 still has 91.6 GiB left for KV cache (our calculation: 164.7 - 155.5 and 247.1 - 155.5). The KV cache guide has the formula for other models.
Compute: dense FP4 goes up, FP8 and BF16 do not#
The B300's compute gain is dense FP4 and attention. Sparse FP4 stays at 18 PFLOPS per GPU, dense FP4 rises from 9 to 14, and FP8, BF16, TF32 and FP32 are unchanged. For attention, NVIDIA doubled the special-function throughput that softmax uses, 10.7 against 5 tera-exponentials per second, and says attention layers run up to 2x faster, with the biggest effect on reasoning models with long contexts. NVIDIA's DGX B300 datasheet sums it up as 1.5x the dense FP4 and 2x the attention performance of a DGX B200.
The losses are just as specific. FP64 drops from 37 to 1.2 TFLOPS per GPU and INT8 from 9 POPS to 307 TOPS. Double-precision simulation and INT8-quantized inference belong on the B200.
Power, cooling and networking#
Power barely moves between the two. A B300 is allowed 1,100 W against the B200's 1,000 W. NVIDIA lists about 14.3 kW at most for a DGX B200 and about 14 kW for a DGX B300, although its February 2026 DGX B300 datasheet gives 14.5 kW at the busbar and 15.1 kW at the power supplies. Both DGX systems are air-cooled and take 10 rack units, so either fits the same room.
Networking doubles. A DGX B300 gives each GPU an 800 Gb/s ConnectX-8 port, against 400 Gb/s ConnectX-7 on a DGX B200, and the host link moves from PCIe Gen5 to Gen6. For multi-node training I would weigh that port as heavily as the extra memory, because between servers every GPU's traffic goes through it. The InfiniBand vs NVLink guide explains the split, and GPU clusters covers the builds.
Which to build for inference and for training#
The deciding question is whether memory, FP4 or attention limits your job.
| Workload | Build on | Why |
|---|---|---|
| Long-context or high-concurrency inference of large models | B300 | 270 GB for weights and KV cache, attention up to 2x faster |
| FP4 (NVFP4) inference | B300 | 14 dense FP4 PFLOPS against 9 |
| FP8 serving that fits in 180 GB with room for KV cache | B200 | Same FP8 Tensor Core throughput |
| FP8 or BF16 training or fine-tuning that fits | B200 | Same FP8 and BF16 throughput per GPU |
| Training limited by memory or by bandwidth between servers | B300 | 1.5x the memory, 800 Gb/s per GPU |
| FP64 simulation or INT8 inference | B200 | 37 against 1.2 TFLOPS of FP64, 9 POPS against 307 TOPS of INT8 |
Software: sm_100 against sm_103#
The B300 needs a newer toolchain than the B200. The B200 runs on CUDA 12.8 or newer with an R570 driver. The B300 needs CUDA 12.9 or newer and an R580 driver, and NVIDIA first lists it in 580.82.07. Both need the open kernel modules on Linux. Kernels built for sm_100a run on a B200 and not on a B300, while the sm_100f family target covers both GPUs from one build. Tools lag too: vLLM's CUDA 12.9 wheels skip compute capability 10.3 and it recommends CUDA 13 for B300, and Axolotl says CUDA 12.8 cannot compile for sm_103a. The driver and CUDA version guide shows how to confirm what a server is running.
Availability and price#
NVIDIA lists both HGX B300 and HGX B200 as shipping, and QuantaCloud builds either to order. Lead times depend on GPU supply and come back in writing with each quote, before you commit.
NVIDIA does not reveal list prices. The closest figure from NVIDIA predates the B300: Jensen Huang told CNBC on 2024-03-19 that a Blackwell GPU would cost $30,000 to $40,000. Server prices vary by configuration, so QuantaCloud quotes both in writing. The B200 page and the B300 page have the full specs, fit tables and build options. The H200 vs B200 and B200 vs H100 comparisons set the B200 against the H200 and the H100. Neither is on demand, so to prototype before the hardware arrives, the largest on-demand GPU is the H200 NVL with 141 GB. Stopping an on-demand instance deletes its disk, so copy results off first.
B300 vs B200 questions#
Is the B300 the same as Blackwell Ultra?
Yes, for servers. NVIDIA calls the GPU "Blackwell Ultra SXM" in HGX B300 and DGX B300. The GB300 superchip uses the same generation with 279 GB per GPU and is the building block of the GB300 NVL72 rack.
Can I test B300 or B200 code before the hardware arrives?
Only partly. The on-demand RTX PRO 6000 Blackwell runs FP4 checkpoints, but it is compute capability 12.0, so it cannot test sm_100 or sm_103 kernels. It is useful for model and serving setup, and the kernels get tested on the delivered servers.
The rule I follow: build on B300 when memory, long context or FP4 inference is the constraint, and on B200 when the job is FP8 or BF16 work that fits in 180 GB per GPU, or needs FP64 or INT8. If you are not sure, send one brief with the model, precision, context length and GPU count, and ask for both configurations. QuantaCloud will quote each in writing with its lead time.
Plan a Blackwell build