Dedicated capacity · built to order

NVIDIA B300

270 GB. Built around your workload.

NVIDIA B300 (Blackwell Ultra) specs from NVIDIA: 270 GB per GPU, dense FP4, power and CUDA needs, plus B300 servers and clusters built to order.

Dedicated hardwareConfiguration scoped with you
A plan for your workloadBuilt to order
Plan dedicated capacity
Configuration and commercial terms agreed in writing.
NVIDIA B300BLACKWELL ULTRA
VRAMVRAMVRAMVRAMVRAMVRAMBLACKWELL ULTRA
Room for the work ahead.270 GB HBM3E
GPU memory
270 GB HBM3E
Memory bandwidth
7.7 TB/s
Architecture
Blackwell Ultra
Deployment
Built to order

The B300 is NVIDIA's Blackwell Ultra GPU: 270 GB of HBM3E per GPU in an HGX B300 server, and about 50% more dense FP4 than the B200. NVIDIA has no standalone B300 product page. The name covers the "Blackwell Ultra SXM" GPUs in HGX B300 and DGX B300 servers, and the same generation sits in the GB300 superchip with 279 GB per GPU. QuantaCloud does not rent the B300 by the hour. We build B300 servers and clusters to order, and the configuration, lead time and terms come back in writing after a capacity brief.

B300 servers built to order#

Each B300 build is ordered against your brief. Tell us the GPU count, network, storage, region, start date and term, and QuantaCloud replies in writing with a configuration, a lead time and commercial terms. We order and build the hardware once you accept. The reserved capacity page lists what to include.

Plan a B300 build

NVIDIA B300 specs#

The B300 keeps the B200's memory bandwidth, FP8 throughput and NVLink, and changes memory size, dense FP4, attention speed, FP64 and INT8. These are NVIDIA's figures, from the Blackwell Ultra datasheet (October 2025) and the HGX and DGX B300 pages as read on 2026-09-28.

SpecPer B300 GPU (HGX B300)HGX B300 system (8 GPUs)
ArchitectureBlackwell Ultra, compute capability 10.3 (sm_103)8x Blackwell Ultra SXM
GPU memory270 GB HBM3E2.1 TB
Memory bandwidth7.7 TB/s62 TB/s
FP4 Tensor Core18 PFLOPS sparse, 14 dense144 PFLOPS sparse, 108 dense
FP8/FP6 Tensor Core9 PFLOPS sparse, 4.5 dense72 PFLOPS sparse, 36 dense
FP16/BF16 Tensor Core4.5 PFLOPS sparse, 2.25 dense36 PFLOPS sparse, 18 dense
INT8 Tensor Core307 TOPS sparse3 POPS sparse
FP641.2 TFLOPS10 TFLOPS
NVLink5th generation, 1.8 TB/s per GPU14.4 TB/s total through NVSwitch
Host linkPCIe Gen6, 256 GB/s
Max powerConfigurable up to 1,100 WDGX B300: about 14 kW
NetworkingConnectX-8, 800 Gb/s per GPU1.6 TB/s
MIGUp to 7 per the datasheet (not yet in NVIDIA's MIG user guide)

Two rows do not multiply out. Eight GPUs at 14 dense FP4 PFLOPS make 112, where the HGX page and the DGX B300 datasheet say 108, and eight at 307 INT8 TOPS make about 2.5 POPS, where the HGX page says 3. I plan on the lower figure in each case.

Against the B200, the GPU itself improves in three ways. Memory rises from 180 to 270 GB per GPU. NVIDIA puts the Blackwell Ultra chip at up to 288 GB and says capacity varies by product, which is why HGX B300 shows 270 GB and GB300 shows 279. Dense FP4 rises from 9 to 14 PFLOPS per GPU while sparse FP4 stays at 18, so the gain lands in dense math, which is what I plan on. Attention gets faster: NVIDIA doubled the special-function throughput that softmax uses and says attention layers run up to 2x faster than on Blackwell. The B300 also gives up most of its FP64, down from 37 to 1.2 TFLOPS per GPU, and most of its INT8, down from 9 POPS to 307 TOPS.

What the B300 is for, and what fits in 270 GB#

The B300 is built for serving large reasoning models at long context, where memory and attention set the pace. The extra 90 GB per GPU goes straight into KV cache or into models that do not fit a B200. The table uses vLLM's default --gpu-memory-utilization of 0.92, which gives the engine 247.1 GiB of each B300 (our calculation: a B300 reports 275,040 MiB, and 275,040 / 1,024 x 0.92 = 247.1), and counts weights and KV cache only, before activations.

Model and precisionWeightsLeft on one B300What that means (our calculation)
Llama-3.3-70B, BF16141.1 GB (131.4 GiB)115.7 GiBAbout 379,000 tokens of BF16 KV cache, or nearly three full 131K contexts. A B200 holds about 109,000 tokens
DeepSeek-V4-Flash, FP4 and FP8About 167 GB on disk (155.5 GiB)91.6 GiBRoom for KV cache and batching. On a B200 the weights leave only 9.2 GiB of the 164.7 GiB the engine gets at the default setting
Kimi K3, MXFP4About 1.56 TB on diskDoes not fit one GPUFits one HGX B300 (2.16 TB), not one HGX B200 (1.44 TB)
Qwen3.8-2.4T, FP82,446 GBDoes not fit one GPUMore than one HGX B300 holds. vLLM's recipe runs it on 16 B300s, which is two servers

The Llama row uses its KV cache of 320 KiB per token at BF16 (our calculation: 2 x 80 layers x 8 KV heads x 128 x 2 bytes). When one server is not enough, the choice is a cluster over InfiniBand or a GB300 NVL72 rack, where 72 GPUs share one NVLink domain.

Servers and clusters we build#

Every B300 build starts from the eight-GPU HGX B300, since that is how NVIDIA ships the B300 for servers. Inside each server, NVSwitch links the eight GPUs at 1.8 TB/s each. Between servers, NVIDIA's DGX B300 gives each GPU its own 800 Gb/s ConnectX-8 port for InfiniBand or RoCE, twice the 400 Gb/s per GPU of a DGX B200. The DGX B300 also uses an air-cooled chassis, available in MGX-compatible and conventional AC-powered versions, fills 10 rack units and draws about 14 kW. An eight-GPU B300 server therefore does not force liquid cooling.

BuildGPUsInterconnectWhere it is described
One single-tenant HGX B300 server8NVLink and NVSwitch inside the serverDedicated GPU servers
B300 cluster16 and up, in steps of 8NVLink inside each server, 800 Gb/s InfiniBand or RoCE between themGPU clusters
Rack-scale Blackwell Ultra72 per NVL72 rackOne 72-GPU NVLink domainGB300 NVL72, discussed before it is quoted

Storage, region and term are part of the quote.

What a B300 costs#

The B300 has no public list price. NVIDIA does not reveal list prices, and the B300 figures that circulate online come from vendor listings, price aggregators and blogs rather than from NVIDIA, so I leave them out. The closest NVIDIA figure is Jensen Huang's 2024 estimate of $30,000 to $40,000 per Blackwell GPU, which predates the B300 and is covered on the B200 page. Server and cluster prices vary with GPU count, fabric, storage and term in any case.

QuantaCloud prices a B300 build as a written quote. The closest dated public figures are for the rack-scale GB300 NVL72, and the reports disagree with each other: the GB300 page has both.

Software readiness for the B300#

The B300 is compute capability 10.3, not 10.0, and your stack has to know it. It needs CUDA 12.9 or newer, the release that added sm_103, and an R580 driver: NVIDIA first lists the B300 in data-center driver 580.82.07, not 580.65.06. On Linux, Blackwell runs only with the open kernel modules. Kernels built for the architecture-specific sm_100a target do not run on a B300, while plain sm_100 code does, under NVIDIA's rule that a binary runs on the same major version with an equal or higher minor. New builds should target sm_103a or the sm_100f family.

Two projects make the point. vLLM recommends CUDA 13 for B300 and GB300, and its CUDA 12.9 wheels skip compute capability 10.3 entirely. Axolotl says CUDA 12.8 cannot compile for sm_103a and asks for PyTorch 2.11 or newer with CUDA 13.0. On a delivered server:

Terminal
nvidia-smi --query-gpu=name,compute_cap,driver_version --format=csv
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"

The first should report 10.3. The architecture list should include sm_100 or sm_103, and sm_100 is enough under the rule above. The driver and CUDA version guide explains each field.

Start on-demand while the build is in progress#

You can prepare on an on-demand GPU while the B300 is on order. The H200 NVL is the largest on-demand card at 141 GB. The RTX PRO 6000 Blackwell runs FP4, but it is compute capability 12.0, so kernels you compile for it will not carry over to the B300's 10.3.

GPUMemoryFromAvailable now
H200 NVL-Not listedNo
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes

Prices checked 6 Oct 2026, 01:40 UTC

On-demand instances are Ubuntu VMs in our US regions, and stopping one deletes its disk, so copy results off first. The deploy guide covers the console steps, and the H200 NVL page has its live configurations.

B300 questions#

Is the B300 the same as Blackwell Ultra?

Yes, in practice. NVIDIA calls the GPU "Blackwell Ultra SXM" inside HGX B300 and DGX B300, and B300 is the name those systems carry. The GB300 superchip pairs the same GPU generation with a Grace CPU, exposes 279 GB per GPU and ships in the GB300 NVL72 rack.

How much memory does the B300 have?

An HGX B300 exposes 270 GB of HBM3E per GPU and 2.1 TB across its eight GPUs. NVIDIA puts the chip's maximum at 288 GB, and the GB300 exposes 279 GB.

Can I rent a B300 by the hour?

Not on QuantaCloud. The B300 is reserved capacity that we build to order. The on-demand catalog on the GPUs page tops out at the H200 NVL with 141 GB.

Should I build on B300 or B200?

Build on B300 for memory-bound and long-context inference. Build on B200 for FP8 or BF16 training that fits in 180 GB per GPU, and for anything that needs FP64 or INT8. The B300 vs B200 comparison puts the numbers side by side.

How long does a B300 build take?

Lead time depends on GPU supply when the order is placed, and QuantaCloud confirms it in writing before you commit.

The rule I follow: build on B300 when the model or its KV cache does not fit comfortably in 180 GB per GPU, or when the job is FP4 inference at long context. If memory is not the constraint and the work is FP8 training, the B200 does the same job, and for FP64 it does a far better one. Send the GPU count, fabric, storage, region, start date and term, and we will come back with a configuration, lead time and terms in writing.

Plan a B300 build

Keep building

Choose your next step.