Rent the H200 NVL when the GPU will be busy serving a large model, and the RTX PRO 6000 when it will not. The H200 NVL has 141 GB of HBM3e at 4.8 TB/s, 1.47 times the RTX PRO 6000's 96 GB and 2.7 to 3.0 times its bandwidth, depending on the RTX PRO 6000's edition. On 2026-09-27 it cost 1.44 times as much per hour, $3.43 against $2.39. So for a model that fits on both, a busy H200 NVL should cost less per token, while an idle hour costs 30% less on the RTX PRO 6000. Two things override that rule: past about 88 GiB only the H200 NVL holds the model on one GPU, and only the RTX PRO 6000 has FP4 Tensor Cores.
RTX PRO 6000 and H200 NVL specs side by side#
The H200 NVL leads on memory, bandwidth and 8 and 16-bit Tensor Core math, and the RTX PRO 6000 on FP4, plain FP32 and video encoding.
| Spec | RTX PRO 6000 Blackwell | H200 NVL |
|---|---|---|
| Architecture (compute capability) | Blackwell (12.0) | Hopper (9.0) |
| GPU memory | 96 GB GDDR7 with ECC | 141 GB HBM3e |
| Memory bandwidth | 1,597 GB/s (Server Edition) or 1,792 GB/s (Workstation and Max-Q) | 4.8 TB/s |
| FP32 | 120 TFLOPS (Server Edition) | 60 TFLOPS |
| BF16 and FP16 Tensor Core (dense) | 503.8 TFLOPS (Workstation Edition) | 835.5 TFLOPS |
| FP8 Tensor Core (dense) | 1,007.6 TFLOPS (Workstation Edition) | 1,670.5 TFLOPS |
| FP4 Tensor Core (dense) | 2,015.2 TFLOPS (Workstation Edition) | Not supported |
| FP64 | Not listed by NVIDIA | 30 TFLOPS, 60 on Tensor Cores |
| NVLink | Not supported | 2- or 4-way bridge, 900 GB/s per GPU |
| Video engines | 4 NVENC, 4 NVDEC, 4 JPEG | 7 NVDEC, 7 JPEG, no encoder listed |
| Host link | PCIe Gen5 x16 | PCIe Gen5, 128 GB/s |
| Max power | 300 to 600 W, by edition | Up to 600 W, configurable |
The two spec sheets need opposite corrections before their Tensor Core rows line up. NVIDIA's H200 page quotes throughput with sparsity, 3,341 FP8 TFLOPS for the NVL, so the table halves it. The RTX PRO 6000's Server Edition page gives only round numbers, 2 PFLOPS of FP8 and 4 of FP4, and those match the sparse rates in NVIDIA's RTX PRO Blackwell architecture whitepaper, which lists dense and sparse rates for the Workstation and Max-Q editions only. The table uses the Workstation Edition's dense rates, the highest of the three editions, and on them the H200 NVL has 1.66 times the dense FP8 and BF16 rates (our calculation: 1,670.5 / 1,007.6). Against the 300 W Max-Q's 877.9 dense FP8 TFLOPS the lead grows to 1.9 times (1,670.5 / 877.9), and the Server Edition, which boosts to 2,430 MHz against the Workstation Edition's 2,617, sits about 7% below the table's rates (2,430 / 2,617 = 0.93). Bandwidth favors the H200 NVL by 2.7 to 3.0 times (4,800 / 1,792 and 4,800 / 1,597). The RTX PRO 6000 has FP4, which Hopper lacks, and on the Server Edition's figure twice the plain FP32 rate (120 / 60).
Prices and cost per token#
The RTX PRO 6000 cost 30% less per GPU-hour on 2026-09-27, and the hourly price is only half of the comparison.
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 6 Oct 2026, 02:10 UTC
On 2026-09-27 a 1x RTX PRO 6000 listed at $2.39 an hour and a 1x H200 NVL at $3.43 (today the console price and the console price). The H200 NVL is cheaper per job whenever it finishes at least 1.44 times as fast (our calculation: 3.43 / 2.39 = 1.44). NVIDIA describes token-by-token generation as memory-bound, and at 2.7 to 3.0 times the bandwidth a busy H200 NVL should generate tokens for about half the cost per token (1.44 / 2.68 = 0.54 and 1.44 / 3.01 = 0.48). The margin is thinner for prefill and training math: at least 1.66 times the dense rate, its lead over the fastest edition, for 1.44 times the price. None of that helps in an hour the GPU spends idle. For a notebook, a model you query now and then or a script you are still debugging, the cheaper hour wins. At the snapshot the RTX PRO 6000 came in 1, 2 and 4-GPU VMs in us-east-1 and us-midwest-4, and the H200 NVL in 1 and 2-GPU VMs in us-east-1, where a 4-GPU VM appeared 13 minutes later. Each configuration is on the RTX PRO 6000 page and the H200 page, and pricing has the billing rules.
The H200 NVL for 70B models, long contexts and heavy traffic#
The H200 NVL is the only one of the two that holds a 70B model at FP8 with its full context. The per-channel FP8 build of Llama-3.3-70B is a 72.7 GB checkpoint, and one full 128k-token sequence adds 42.9 GB of BF16 KV cache, 115.6 GB in all (107.7 GiB), inside the H200 NVL's 129.2 GiB budget. On the RTX PRO 6000 the same checkpoint leaves about 20.3 GiB, roughly 66,000 tokens of BF16 KV cache shared by every request, or about 133,000 with --kv-cache-dtype fp8 (our calculation at 320 and 160 KiB per token). For long documents or agent transcripts on a 70B model, that difference decides it.
Heavy traffic is the second case. Throughput under load depends on how many sequences fit in the KV cache and how fast the GPU reads memory, and the H200 NVL leads on both. gpt-oss-120b leaves room for about 15 full 131k-token sequences on the H200 NVL and about 6 on the RTX PRO 6000 (our calculation: (129.2 - 60.8) / 4.5 = 15.2 and (87.9 - 60.8) / 4.5 = 6.0). A shared endpoint that stays busy is where the lower cost per token shows up.
A pair of cards is the third case, if the bridge is there. The H200 NVL supports a 2- or 4-way NVLink bridge at 900 GB/s per GPU, and QuantaCloud's API flags its multi-GPU offers as NVLink, while the RTX PRO 6000 has no NVLink at all. Llama-3.3-70B in BF16, 141.1 GB of weights, fits on 2x H200 NVL with about 127 GiB left for KV cache, and on 2x RTX PRO 6000 with 44.5 GiB left and the model split over PCIe (our calculation: 2 x 129.17 - 131.42 and 2 x 87.95 - 131.42). Run nvidia-smi topo -m on the 2x H200 NVL VM before you plan on tensor parallelism across it. The H200 NVL is also the card for double precision, at 30 TFLOPS of FP64, where NVIDIA lists no FP64 figure for the RTX PRO 6000.
The RTX PRO 6000 for FP4, video and idle hours#
FP4 is the RTX PRO 6000's clearest win. vLLM runs NVFP4 checkpoints natively on Blackwell and falls back to weight-only 4-bit kernels on the H200 NVL. NVIDIA rates the Workstation Edition at 2,015.2 dense FP4 TFLOPS, 1.21 times the H200 NVL's dense FP8 rate (our calculation: 2,015.2 / 1,670.5), and the Max-Q at 1,755.7, 1.05 times, at 30% less per hour either way. FP4 weights also take about half the memory of FP8. The trade is accuracy, so compare an NVFP4 build with its FP8 original on your own prompts before a production model moves to it.
Image and video generation is the second. ComfyUI computes in NVFP4 only on compute capability 10 and newer with a cu130 PyTorch build, which leaves out the H200 NVL at 9.0. BFL's NVFP4 file for FLUX.1 [dev] is 9.2 GB, and FLUX.2 [dev]'s fp8 transformer and fp8 text encoder, 53.5 GB together, fit on either card. The RTX PRO 6000 also has four NVENC encoders for the finished video, while NVIDIA's spec table for the H200 NVL lists only decoders. The FLUX in ComfyUI guide lists the files with the licence terms for each.
Hours you do not fill are the third. A notebook you are exploring, a ComfyUI graph you are still building or an assistant a few people use leaves the GPU idle between requests, and you pay for the hour either way. On 2026-09-27 the RTX PRO 6000 hour was $1.04 cheaper, and it came in two regions and in 4-GPU VMs with 384 GB in total. The Server Edition's 120 TFLOPS of FP32, twice the H200 NVL's 60, also helps code that does not run on Tensor Cores.
What fits in 96 GB and in 141 GB#
The 41.2 GiB between the two working budgets decides most of these rows: 87.9 GiB on the RTX PRO 6000 and 129.2 GiB on the H200 NVL, with vLLM reserving its default 92% of GPU memory (our calculation from the 97,887 and 143,771 MiB that nvidia-smi reports).
| Workload | Memory it needs | Source | RTX PRO 6000, 96 GB | H200 NVL, 141 GB |
|---|---|---|---|---|
| gpt-oss-120b | About 65 GB, plus 4.5 GiB per full 131k-token sequence | OpenAI, our calculation | Fits, 6.0 full sequences | Fits, 15.2 full sequences |
| Qwen3-32B, BF16, 32k context | 65.5 GB, plus 8 GiB per sequence | Our calculation | Fits, 3.4 sequences | Fits, 8.5 sequences |
| Llama-3.3-70B, per-channel FP8 | 72.7 GB checkpoint | Hugging Face (RedHatAI) | Fits, about 66,000 tokens of KV cache | Fits, about 200,000 tokens |
| Llama-3.3-70B, FP8, one full 128k sequence | 72.7 + 42.9 = 115.6 GB | Our calculation | Does not fit | Fits |
| Llama-3.3-70B, BF16 | 141.1 GB of weights | Our calculation | 2 GPUs, 44.5 GiB left, over PCIe | 2 GPUs, about 127 GiB left |
| FLUX.1 [dev], NVFP4 | 9.2 GB file | Black Forest Labs | Fits, FP4 compute | Fits, no FP4 compute |
| FLUX.2 [dev], fp8 transformer plus fp8 text encoder | 35.5 GB + 18.0 GB | Comfy-Org files | Fits | Fits |
| QLoRA, gpt-oss-120b | 65 GB | Unsloth | Fits | Fits |
| 16-bit LoRA, 70B model | 164 GB | Unsloth | 2 GPUs | 2 GPUs |
The KV cache figures are BF16, before activations and CUDA graph memory, so treat them as ceilings. The VRAM guide runs the same arithmetic for other models, the KV cache guide works through the per-token formula, and gpt-oss GPU requirements covers the 120b model in detail.
FAQ#
Is one H200 NVL better than two RTX PRO 6000s?
For a model that fits in 141 GB, usually. On 2026-09-27 a 2x RTX PRO 6000 VM listed at $4.79 an hour, 1.40 times the price of one H200 NVL, for 192 GB against 141. Llama-3.3-70B at FP8 with a full 128k context, 115.6 GB, fits on either. The pair splits the model over PCIe, where vLLM recommends pipeline parallelism, while the single H200 NVL keeps it on one GPU with more bandwidth than both RTX PRO 6000 cards combined (4,800 GB/s against at most 2 x 1,792 = 3,584). The pair makes sense for two independent jobs, or for a model that needs more than the H200 NVL's 129.2 GiB budget.
Can the H200 NVL run FP4 models?
It loads them and computes them in 16-bit. FP4 Tensor Cores are Blackwell-only, so vLLM runs NVFP4 checkpoints on the H200 NVL as weight-only 4-bit, and ComfyUI's NVFP4 path needs compute capability 10 or newer. The memory saving carries over to the H200 NVL, and the FP4 speed does not.
Do the 2x VMs link their GPUs with NVLink?
The RTX PRO 6000 VMs never do, because NVIDIA lists NVLink as not supported on that card. The H200 NVL supports a bridge, and QuantaCloud's API flags those offers as NVLink, but I only treat it as settled when nvidia-smi topo -m shows NV entries between the GPUs. The InfiniBand vs NVLink guide explains what the link carries.
What happens to a downloaded checkpoint when I stop?
It goes with the VM's local disk, 725 GB on a 1x RTX PRO 6000 and 750 GB on a 1x H200 NVL on 2026-09-27. Stopping either instance terminates it, and there are no volumes or snapshots, so a 72.7 GB checkpoint you downloaded has to be downloaded again on the next VM. What you do get back is the unused part of the current hour, refunded by the second (pricing).
The one thing I check first is how busy the GPU will be. If the model and its KV cache need more than about 88 GiB, or the GPU will serve traffic most of the hour, rent the H200 NVL. If the job fits in 88 GiB and the GPU will sit idle between runs, or you want NVFP4 in vLLM or ComfyUI, rent the RTX PRO 6000. If you cannot tell, time your own prompts on a 1x of each: at the 2026-09-27 prices, 20 minutes costs about $0.80 on the RTX PRO 6000 and $1.14 on the H200 NVL (our calculation: a third of the hourly price), since billing runs from launch to stop. RTX PRO 6000 vs H100 and A100 vs H200 cover the 80 GB options, and H200 vs B200 covers the step past 141 GB.
Launch an H200 NVL Launch an RTX PRO 6000