GPU comparison

RTX PRO 6000 Blackwell vs H100 PCIe

H100 PCIe vs RTX PRO 6000 Blackwell for LLM inference: 80 GB of faster HBM2e against 96 GB with FP4, at a similar hourly price. Tokens per dollar on paper.

Faiz Ahmed11 min read
NVIDIA · Blackwell

RTX PRO 6000 Blackwell

Cloud GPU with full SSH access

96 GBGDDR7 memory
Bandwidth
1,597–1,792 GB/s by edition
FP8 compute
Native FP8
From / GPU-hrSee console
Explore RTX PRO 6000 Blackwell
NVIDIA · Hopper

H100 PCIe

Cloud GPU with full SSH access

80 GBHBM2e memory
Bandwidth
2.0 TB/s
FP8 compute
Native FP8
From / GPU-hr$2.59
Launch H100 PCIe

Prices checked

For inference, the H100 PCIe is the better value when the model and its KV cache fit in about 73 GiB, and the RTX PRO 6000 when they need more or when you serve NVFP4 checkpoints. The two cost nearly the same: on 2026-09-27 a 1x H100 PCIe listed at $2.59 an hour and a 1x RTX PRO 6000 at $2.39, 8% less. For that 8% the H100 reads memory 1.12 to 1.25 times as fast, depending on the RTX PRO 6000's edition, and computes dense FP8 at least 1.5 times as fast, which is its margin over the fastest RTX PRO 6000 edition. The RTX PRO 6000 answers with 16 GB more memory, FP4 Tensor Cores and, on the Server Edition's figure, 2.3 times the plain FP32 rate.

RTX PRO 6000 and H100 PCIe specs side by side#

The H100 PCIe leads on bandwidth and 8 and 16-bit Tensor Core math, and the RTX PRO 6000 on memory, FP4 and FP32. Both are PCIe cards, so the table compares card with card.

SpecRTX PRO 6000 BlackwellH100 PCIe
Architecture (compute capability)Blackwell (12.0)Hopper (9.0)
GPU memory96 GB GDDR7 with ECC80 GB HBM2e
Memory bandwidth1,597 GB/s (Server Edition) or 1,792 GB/s (Workstation and Max-Q)2,000 GB/s
FP32120 TFLOPS (Server Edition)51.2 TFLOPS
BF16 and FP16 Tensor Core (dense)503.8 TFLOPS (Workstation Edition)756 TFLOPS
FP8 Tensor Core (dense)1,007.6 TFLOPS (Workstation Edition)1,513 TFLOPS
FP4 Tensor Core (dense)2,015.2 TFLOPS (Workstation Edition)Not supported
FP64Not listed by NVIDIA25.6 TFLOPS, 51.2 on Tensor Cores
NVLinkNot supportedBridge to one adjacent card, 600 GB/s
Multi-Instance GPUUp to 4 x 24 GBUp to 7 x 10 GB
Video encoders (NVENC)4None
Host linkPCIe Gen5 x16PCIe Gen5 x16, 128 GB/s
Max power300 to 600 W, by edition350 W

The H100 PCIe figures come from NVIDIA's H100 whitepaper and product brief, and the RTX PRO 6000's memory, bandwidth, FP32 and power from its Server Edition page and datasheet. For the RTX PRO 6000's Tensor Core rows I use NVIDIA's RTX PRO Blackwell architecture whitepaper, the one NVIDIA document with dense rates for the card, and it covers only the 600 W Workstation Edition and the 300 W Max-Q. The Server Edition page rounds to 2 PFLOPS of FP8 and 1 PFLOP of BF16, which match the whitepaper's sparse rates, so setting them against the H100's dense 1,513 and 756 would double the RTX PRO 6000's side. The edition changes the size of the H100's lead, not who has it. The H100 has 1.5 times the Workstation Edition's dense FP8 and BF16 rates (our calculation: 1,513 / 1,007.6 and 756 / 503.8), 1.7 times the Max-Q's (1,513 / 877.9), and about 1.6 times a Server Edition's, which boosts to 2,430 MHz against the Workstation Edition's 2,617 (1.50 / 0.93). It also reads memory 1.12 to 1.25 times as fast (2,000 / 1,792 and 2,000 / 1,597). The RTX PRO 6000 has 1.2 times the memory, FP4, and on the Server Edition's figure 2.3 times the FP32 rate (120 / 51.2).

Prices and value per token#

The RTX PRO 6000 cost 8% less per GPU-hour on 2026-09-27, a smaller gap than any of the H100's paper advantages.

GPUMemoryFromAvailable now
RTX PRO 6000 Blackwell-Not listedNo
H100 PCIe80 GB$2.59/GPU-hrYes

Prices checked 5 Oct 2026, 17:50 UTC

On 2026-09-27 a 1x RTX PRO 6000 listed at $2.39 an hour and a 1x H100 PCIe at $2.59 (today the console price and $2.59/GPU-hr). The H100 costs less per token whenever it runs at least 1.08 times as fast (our calculation: 2.59 / 2.39 = 1.08). Its 1.12 to 1.25 times the bandwidth clears that bar for token generation, which NVIDIA describes as memory-bound, and its 1.5 to 1.7 times the dense FP8 rate, by RTX PRO 6000 edition, clears it for prefill. So on a model that fits both cards with room for its KV cache, the H100 PCIe should deliver more tokens per dollar. The RTX PRO 6000 wins where that condition breaks: the model needs more than 73 GiB, the traffic needs more KV cache than the H100 has left, or the checkpoint is NVFP4. That day the RTX PRO 6000 came in 1, 2 and 4-GPU VMs in us-east-1 and us-midwest-4 and the H100 PCIe in 1 and 2-GPU VMs in us-midwest-2, and the RTX PRO 6000 page and the H100 page list the vCPUs, RAM and disk of each. The pricing page explains how an hour is billed.

Under 73 GiB, the H100 PCIe gives more tokens per dollar#

A model that fits with room to spare is the H100's case. Qwen3-32B at FP8 is a 34.3 GB checkpoint, per-channel or Qwen's own block-scaled build, and vLLM 0.30.0 runs both formats with FP8 math on Hopper and Blackwell alike. On the H100 it leaves room for 5.2 full 32k-token conversations at once ((73.3 - 32.0) / 8), and each token comes out faster than on the RTX PRO 6000 for 8% more per hour.

gpt-oss-120b is the same story with less room. OpenAI sized it for a single 80 GB GPU, and on the H100 its 65 GB of weights leave space for 2.8 full 131k-token sequences, or many more short ones. For one user or short conversations that is enough, and the faster memory makes each response quicker. gpt-oss GPU requirements covers the model in detail.

Software is the third reason. Hopper has been supported since CUDA 11.8, so every current PyTorch and vLLM build has kernels for it, and FlashAttention-3 lists the H100 as a requirement. Blackwell needs CUDA 12.8 or newer builds, and an image with a PyTorch build for CUDA 12.6 fails on the RTX PRO 6000 with "no kernel image is available for execution on the device". The H100 PCIe also supports an NVLink bridge to one neighboring card and runs FP64 at 25.6 TFLOPS, while NVIDIA lists NVLink as not supported on the RTX PRO 6000 and gives it no FP64 figure.

Over 73 GiB, or with NVFP4, the RTX PRO 6000#

KV cache room is the RTX PRO 6000's first advantage, and a 70B model at FP8 shows it most plainly. Both cards run the per-channel FP8 build of Llama-3.3-70B with FP8 math, so the only difference is what is left after its 72.7 GB of weights: about 5.6 GiB of the H100's 73.3 GiB budget, against about 20.3 GiB on the RTX PRO 6000, roughly 66,000 tokens of BF16 KV cache or 133,000 with --kv-cache-dtype fp8 (our calculation at 320 and 160 KiB per token). On the H100 that checkpoint loads and then has room for about 18,000 tokens, one long conversation at most.

Long contexts across many users are the second. KV cache capacity caps how many requests a server holds at once, and gpt-oss-120b leaves room for 6.0 full-length sequences on the RTX PRO 6000 against 2.8 on the H100. For Qwen3-32B at FP8 the count is 7.0 full 32k-token conversations against 5.2 (our calculation: (87.9 - 32.0) / 8). For a shared endpoint with long documents or agent transcripts, the extra 14.7 GiB of working budget becomes more concurrent users (87.95 - 73.28).

NVFP4 is the third. vLLM runs NVFP4 checkpoints natively on Blackwell and weight-only on the H100, which has no FP4 Tensor Cores. NVIDIA rates the Workstation Edition at 2,015.2 dense FP4 TFLOPS, 1.33 times the H100's dense FP8 rate (our calculation: 2,015.2 / 1,513), and the Max-Q at 1,755.7, 1.16 times, with weights about half the size of FP8. NVFP4 gives up some accuracy for that, so run your own evaluation prompts against the FP8 build before you switch a production model.

It is also the better card when image or video work shares the GPU with an LLM. ComfyUI computes in NVFP4 only on compute capability 10 and newer, and the RTX PRO 6000 has four NVENC encoders, where NVIDIA's H100 whitepaper notes the H100 has none. On 2026-09-27 it also came in 4-GPU VMs with 384 GB in total, while the H100 PCIe stopped at two.

What fits in 80 GB and in 96 GB#

Serving room is the useful way to read this table: 73.3 GiB of working budget on the H100 PCIe against 87.9 GiB on the RTX PRO 6000, at vLLM's default of reserving 92% of GPU memory (our calculation from the 81,559 and 97,887 MiB that nvidia-smi reports).

WorkloadMemory it needsSourceH100 PCIe, 80 GBRTX PRO 6000, 96 GB
Qwen3-32B, FP8, 32k conversations34.3 GB checkpoint, plus 8 GiB per conversationHugging Face, our calculationFits, 5.2 at onceFits, 7.0 at once
gpt-oss-120bAbout 65 GB, plus 4.5 GiB per full 131k-token sequenceOpenAI, our calculationFits, 2.8 full sequencesFits, 6.0 full sequences
Gemma 4 31B, BF1662.5 GB of weightsGoogle, our calculationFits, per GoogleFits
Llama-3.3-70B, 4-bit AWQ39.8 GB checkpointHugging FaceFits, about 119,000 tokens of KV cacheFits, about 167,000 tokens
Qwen3-32B, BF16, one 32k sequence69.0 GiBOur calculationJust inside the budget, 4.3 GiB spareFits, about 19 GiB spare
Llama-3.3-70B, per-channel FP872.7 GB checkpointHugging Face (RedHatAI)Loads, about 5.6 GiB left, about 18,000 tokensFits, about 66,000 tokens of KV cache
Llama-3.3-70B, FP8, one full 128k sequence72.7 + 42.9 = 115.6 GBOur calculationDoes not fitDoes not fit, use an H200 NVL
FLUX.1 [dev], NVFP49.2 GB fileBlack Forest LabsFits, no FP4 computeFits, FP4 compute

The per-token KV figures behind these rows come from each model's config, and the KV cache guide shows how to work them out. For a model not listed, start from the VRAM guide.

FAQ#

Does the RTX PRO 6000's edition change the result?

Not the winner, only the margin. Against the H100's 2,000 GB/s, the Server Edition's 1,597 GB/s gives the H100 a 1.25 times bandwidth lead, and the 1,792 GB/s of the Workstation and Max-Q editions a 1.12 times lead, both above the 1.08 price ratio. nvidia-smi -q on the VM names the edition, since the catalog lists only RTX PRO 6000 Blackwell.

What if the model needs more than 96 GB?

Then compare one bigger card with a pair. On 2026-09-27 a 2x RTX PRO 6000 VM listed at $4.79 an hour and a 2x H100 PCIe VM at $5.18, while one H200 NVL with 141 GB cost $3.43. For Llama-3.3-70B at FP8 with a full 128k context, 115.6 GB, the single H200 NVL is the cheaper way to hold it and keeps the model on one GPU. A pair of RTX PRO 6000 cards always splits over PCIe, where vLLM recommends pipeline parallelism, and whether a pair of H100s has an NVLink bridge only nvidia-smi topo -m on the VM can tell. RTX PRO 6000 vs H200 covers that step.

Do the model weights survive a stop?

No, they are deleted with the VM's local disk, 1,250 GB on a 1x H100 PCIe and 725 GB on a 1x RTX PRO 6000 on 2026-09-27. Stopping terminates the instance, and there are no volumes or snapshots, so a 72.7 GB FP8 checkpoint has to be downloaded again on the next VM, and fast Hugging Face downloads shortens that wait. Your balance gets back the unused seconds of the current hour, as the pricing page explains.


For inference, my test for this pair is the working budget. Add up the weights and the KV cache your traffic needs. Under about 73 GiB, rent the H100 PCIe: it reads memory faster and computes FP8 faster for 8% more per hour. Between 73 and 88 GiB, or for NVFP4 checkpoints, rent the RTX PRO 6000. Past 88 GiB, go to the H200 NVL rather than splitting the model. Then check cost per million tokens on your own prompts: the hourly price divided by (tokens per second x 3,600), times 1,000,000. L40S vs H100 and RTX PRO 6000 vs A100 cover the cheaper cards on each side.

Launch an H100 PCIe Launch an RTX PRO 6000

Keep building

Choose your next step.