The B300 is NVIDIA's Blackwell Ultra GPU: 270 GB of HBM3E per GPU in an HGX B300 server, and about 50% more dense FP4 than the B200. NVIDIA has no standalone B300 product page. The name covers the "Blackwell Ultra SXM" GPUs in HGX B300 and DGX B300 servers, and the same generation sits in the GB300 superchip with 279 GB per GPU. QuantaCloud does not rent the B300 by the hour. We build B300 servers and clusters to order, and the configuration, lead time and terms come back in writing after a capacity brief.
B300 servers built to order#
Each B300 build is ordered against your brief. Tell us the GPU count, network, storage, region, start date and term, and QuantaCloud replies in writing with a configuration, a lead time and commercial terms. We order and build the hardware once you accept. The reserved capacity page lists what to include.
Plan a B300 buildNVIDIA B300 specs#
The B300 keeps the B200's memory bandwidth, FP8 throughput and NVLink, and changes memory size, dense FP4, attention speed, FP64 and INT8. These are NVIDIA's figures, from the Blackwell Ultra datasheet (October 2025) and the HGX and DGX B300 pages as read on 2026-09-28.
| Spec | Per B300 GPU (HGX B300) | HGX B300 system (8 GPUs) |
|---|---|---|
| Architecture | Blackwell Ultra, compute capability 10.3 (sm_103) | 8x Blackwell Ultra SXM |
| GPU memory | 270 GB HBM3E | 2.1 TB |
| Memory bandwidth | 7.7 TB/s | 62 TB/s |
| FP4 Tensor Core | 18 PFLOPS sparse, 14 dense | 144 PFLOPS sparse, 108 dense |
| FP8/FP6 Tensor Core | 9 PFLOPS sparse, 4.5 dense | 72 PFLOPS sparse, 36 dense |
| FP16/BF16 Tensor Core | 4.5 PFLOPS sparse, 2.25 dense | 36 PFLOPS sparse, 18 dense |
| INT8 Tensor Core | 307 TOPS sparse | 3 POPS sparse |
| FP64 | 1.2 TFLOPS | 10 TFLOPS |
| NVLink | 5th generation, 1.8 TB/s per GPU | 14.4 TB/s total through NVSwitch |
| Host link | PCIe Gen6, 256 GB/s | |
| Max power | Configurable up to 1,100 W | DGX B300: about 14 kW |
| Networking | ConnectX-8, 800 Gb/s per GPU | 1.6 TB/s |
| MIG | Up to 7 per the datasheet (not yet in NVIDIA's MIG user guide) |
Two rows do not multiply out. Eight GPUs at 14 dense FP4 PFLOPS make 112, where the HGX page and the DGX B300 datasheet say 108, and eight at 307 INT8 TOPS make about 2.5 POPS, where the HGX page says 3. I plan on the lower figure in each case.
Against the B200, the GPU itself improves in three ways. Memory rises from 180 to 270 GB per GPU. NVIDIA puts the Blackwell Ultra chip at up to 288 GB and says capacity varies by product, which is why HGX B300 shows 270 GB and GB300 shows 279. Dense FP4 rises from 9 to 14 PFLOPS per GPU while sparse FP4 stays at 18, so the gain lands in dense math, which is what I plan on. Attention gets faster: NVIDIA doubled the special-function throughput that softmax uses and says attention layers run up to 2x faster than on Blackwell. The B300 also gives up most of its FP64, down from 37 to 1.2 TFLOPS per GPU, and most of its INT8, down from 9 POPS to 307 TOPS.
What the B300 is for, and what fits in 270 GB#
The B300 is built for serving large reasoning models at long context, where memory and attention set the pace. The extra 90 GB per GPU goes straight into KV cache or into models that do not fit a B200. The table uses vLLM's default --gpu-memory-utilization of 0.92, which gives the engine 247.1 GiB of each B300 (our calculation: a B300 reports 275,040 MiB, and 275,040 / 1,024 x 0.92 = 247.1), and counts weights and KV cache only, before activations.
| Model and precision | Weights | Left on one B300 | What that means (our calculation) |
|---|---|---|---|
| Llama-3.3-70B, BF16 | 141.1 GB (131.4 GiB) | 115.7 GiB | About 379,000 tokens of BF16 KV cache, or nearly three full 131K contexts. A B200 holds about 109,000 tokens |
| DeepSeek-V4-Flash, FP4 and FP8 | About 167 GB on disk (155.5 GiB) | 91.6 GiB | Room for KV cache and batching. On a B200 the weights leave only 9.2 GiB of the 164.7 GiB the engine gets at the default setting |
| Kimi K3, MXFP4 | About 1.56 TB on disk | Does not fit one GPU | Fits one HGX B300 (2.16 TB), not one HGX B200 (1.44 TB) |
| Qwen3.8-2.4T, FP8 | 2,446 GB | Does not fit one GPU | More than one HGX B300 holds. vLLM's recipe runs it on 16 B300s, which is two servers |
The Llama row uses its KV cache of 320 KiB per token at BF16 (our calculation: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). When one server is not enough, the choice is a cluster over InfiniBand or a GB300 NVL72 rack, where 72 GPUs share one NVLink domain.
Servers and clusters we build#
Every B300 build starts from the eight-GPU HGX B300, since that is how NVIDIA ships the B300 for servers. Inside each server, NVSwitch links the eight GPUs at 1.8 TB/s each. Between servers, NVIDIA's DGX B300 gives each GPU its own 800 Gb/s ConnectX-8 port for InfiniBand or RoCE, twice the 400 Gb/s per GPU of a DGX B200. The DGX B300 also uses an air-cooled chassis, available in MGX-compatible and conventional AC-powered versions, fills 10 rack units and draws about 14 kW. An eight-GPU B300 server therefore does not force liquid cooling.
| Build | GPUs | Interconnect | Where it is described |
|---|---|---|---|
| One single-tenant HGX B300 server | 8 | NVLink and NVSwitch inside the server | Dedicated GPU servers |
| B300 cluster | 16 and up, in steps of 8 | NVLink inside each server, 800 Gb/s InfiniBand or RoCE between them | GPU clusters |
| Rack-scale Blackwell Ultra | 72 per NVL72 rack | One 72-GPU NVLink domain | GB300 NVL72, discussed before it is quoted |
Storage, region and term are part of the quote.
What a B300 costs#
The B300 has no public list price. NVIDIA does not reveal list prices, and the B300 figures that circulate online come from vendor listings, price aggregators and blogs rather than from NVIDIA, so I leave them out. The closest NVIDIA figure is Jensen Huang's 2024 estimate of $30,000 to $40,000 per Blackwell GPU, which predates the B300 and is covered on the B200 page. Server and cluster prices vary with GPU count, fabric, storage and term in any case.
QuantaCloud prices a B300 build as a written quote. The closest dated public figures are for the rack-scale GB300 NVL72, and the reports disagree with each other: the GB300 page has both.
Software readiness for the B300#
The B300 is compute capability 10.3, not 10.0, and your stack has to know it. It needs CUDA 12.9 or newer, the release that added sm_103, and an R580 driver: NVIDIA first lists the B300 in data-center driver 580.82.07, not 580.65.06. On Linux, Blackwell runs only with the open kernel modules. Kernels built for the architecture-specific sm_100a target do not run on a B300, while plain sm_100 code does, under NVIDIA's rule that a binary runs on the same major version with an equal or higher minor. New builds should target sm_103a or the sm_100f family.
Two projects make the point. vLLM recommends CUDA 13 for B300 and GB300, and its CUDA 12.9 wheels skip compute capability 10.3 entirely. Axolotl says CUDA 12.8 cannot compile for sm_103a and asks for PyTorch 2.11 or newer with CUDA 13.0. On a delivered server:
nvidia-smi --query-gpu=name,compute_cap,driver_version --format=csv
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"
The first should report 10.3. The architecture list should include sm_100 or sm_103, and sm_100 is enough under the rule above. The driver and CUDA version guide explains each field.
Start on-demand while the build is in progress#
You can prepare on an on-demand GPU while the B300 is on order. The H200 NVL is the largest on-demand card at 141 GB. The RTX PRO 6000 Blackwell runs FP4, but it is compute capability 12.0, so kernels you compile for it will not carry over to the B300's 10.3.
| GPU | Memory | From | Available now |
|---|---|---|---|
| H200 NVL | - | Not listed | No |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
Prices checked 6 Oct 2026, 01:40 UTC
On-demand instances are Ubuntu VMs in our US regions, and stopping one deletes its disk, so copy results off first. The deploy guide covers the console steps, and the H200 NVL page has its live configurations.
B300 questions#
Is the B300 the same as Blackwell Ultra?
Yes, in practice. NVIDIA calls the GPU "Blackwell Ultra SXM" inside HGX B300 and DGX B300, and B300 is the name those systems carry. The GB300 superchip pairs the same GPU generation with a Grace CPU, exposes 279 GB per GPU and ships in the GB300 NVL72 rack.
How much memory does the B300 have?
An HGX B300 exposes 270 GB of HBM3E per GPU and 2.1 TB across its eight GPUs. NVIDIA puts the chip's maximum at 288 GB, and the GB300 exposes 279 GB.
Can I rent a B300 by the hour?
Not on QuantaCloud. The B300 is reserved capacity that we build to order. The on-demand catalog on the GPUs page tops out at the H200 NVL with 141 GB.
Should I build on B300 or B200?
Build on B300 for memory-bound and long-context inference. Build on B200 for FP8 or BF16 training that fits in 180 GB per GPU, and for anything that needs FP64 or INT8. The B300 vs B200 comparison puts the numbers side by side.
How long does a B300 build take?
Lead time depends on GPU supply when the order is placed, and QuantaCloud confirms it in writing before you commit.
The rule I follow: build on B300 when the model or its KV cache does not fit comfortably in 180 GB per GPU, or when the job is FP4 inference at long context. If memory is not the constraint and the work is FP8 training, the B200 does the same job, and for FP64 it does a far better one. Send the GPU count, fabric, storage, region, start date and term, and we will come back with a configuration, lead time and terms in writing.
Plan a B300 build