The RTX PRO 6000 Blackwell is the GPU to rent when a job needs 96 GB on one card. On QuantaCloud, RTX PRO 6000 cloud VMs run from $2.39/GPU-hr, with 1, 2 or 4 cards per VM, and on 2026-09-27 it was the only GPU in the on-demand catalog with FP4 Tensor Cores. The catch is software: Blackwell needs CUDA 12.8 or newer builds, and many older containers fail on it.
Each configuration is a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu, and its local disk is part of the hourly price. On 2026-09-27 the 1x came with 16 vCPUs, 144 GB of RAM and 725 GB of disk, and the 4x with 60 vCPUs, 576 GB of RAM and 2,900 GB, in us-east-1 (Virginia) and us-midwest-4. The first hour is charged at launch and the unused seconds of the current hour are refunded when you stop. The pricing page has the rules, and the GPU catalog has every other card.
Which RTX PRO 6000 this is#
The edition matters more on this GPU than on any other in the catalog. NVIDIA sells three RTX PRO 6000 Blackwell editions with the same 96 GB and different power and bandwidth:
| Edition | Cooling | Power | Memory bandwidth |
|---|---|---|---|
| Server Edition | Passive, for server airflow | Up to 600 W, configurable (600 W or 450 W mode, 300 W minimum) | 1,597 GB/s |
| Workstation Edition | Flow-through fan | 600 W | 1,792 GB/s |
| Max-Q Workstation Edition | Blower fan | 300 W | 1,792 GB/s |
The catalog lists the card only as "RTX PRO 6000 Blackwell", and nvidia-smi -q on any VM you launch shows which edition it is.
RTX PRO 6000 Blackwell specs#
The RTX PRO 6000 Blackwell is a 96 GB card with FP4 Tensor Cores and no NVLink. The figures below are from NVIDIA's Server Edition product page and datasheet, with the dense Tensor rates worked out from NVIDIA's architecture whitepaper. All three editions share the 96 GB, compute capability 12.0 and the four-instance MIG limit.
| Spec | RTX PRO 6000 Blackwell Server Edition |
|---|---|
| Architecture | NVIDIA Blackwell (GB202), compute capability 12.0 |
| GPU memory | 96 GB GDDR7 with ECC, 512-bit |
| Memory bandwidth | 1,597 GB/s (Workstation and Max-Q: 1,792 GB/s) |
| CUDA cores | 24,064 |
| Tensor Cores | 752, fifth generation, with FP8 and FP4 |
| RT Cores | 188, fourth generation |
| FP32 | 120 TFLOPS |
| FP8 and FP4 Tensor | 2 PFLOPS and 4 PFLOPS with sparsity. Dense, about 936 and 1,871 TFLOPS at the Server Edition's clock (our calculation), and 1,007.6 and 2,015.2 on the Workstation Edition |
| NVLink | Not supported |
| MIG | Up to 4 instances of 24 GB |
| System interface | PCIe 5.0 x16 |
| Power | Up to 600 W, configurable |
| Video engines | 4 NVENC, 4 NVDEC, 4 JPEG |
Do not confuse it with the RTX 6000 Ada (Ada, compute capability 8.9) or the RTX A6000 (Ampere, 8.6). Both of those are 48 GB cards, half the memory of the PRO 6000 Blackwell. For the RTX 5090, see RTX PRO 6000 vs RTX 5090.
Software that runs on Blackwell#
The core problem with Blackwell is software, not hardware. It needs CUDA 12.8 or newer builds, and PyTorch builds for CUDA 12.6 or older have no kernels for it. When a stack is too old, you see one of these errors:
... with CUDA capability sm_120 is not compatible with the current PyTorch installation.
... no kernel image is available for execution on the device
Two commands tell you where you stand. The first prints the driver version, and the second must list sm_120:
nvidia-smi
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"
The driver decides which builds run. This is the table I work from:
| Driver on the VM | PyTorch | vLLM 0.30.0 Docker image | ComfyUI v0.37 |
|---|---|---|---|
| 580 or newer | pip install torch works: on Linux it installs CUDA 13.0 wheels | vllm/vllm-openai:v0.30.0 | Runs. It requires a cu130 PyTorch build |
| 570 to 575 | Plain pip install torch installs CUDA 13.0 wheels this driver cannot run. Pin a cu129 build (Linux, PyTorch 2.13 and older) or cu128 (2.11 and older) | vllm/vllm-openai:v0.30.0-cu129 | Needs a 580 driver first |
| Older than 570 | No Blackwell support | Not supported | Not supported |
A cu129 pin looks like pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu129. The driver and CUDA version guide explains the checks, and installing vLLM and the vLLM Docker guide show where the tag goes.
What 96 GB lets you run#
The rule I follow is the same as on any card: weights plus KV cache under 92% of memory, which is about 88 GiB here (our calculation: the card reports 97,887 MiB, and 97,887 / 1,024 x 0.92 = 87.9 GiB, vLLM's default).
| Workload | Memory it needs | On one RTX PRO 6000 | Basis |
|---|---|---|---|
| gpt-oss-120b | About 65 GB (60.8 GiB) on disk, plus 4.5 GiB of KV per full 131k-token sequence | Yes, with KV room for about six full-length sequences at once | Published sizes, our calculation |
| Qwen3-32B in BF16, one 32k sequence | About 69 GiB | Yes, with about 19 GiB to spare | Our calculation |
| Llama-3.3-70B in FP8 | 72.7 GB (67.7 GiB) checkpoint, leaving about 20.3 GiB for KV | Tight: about 66k tokens of BF16 KV, or twice that with FP8 KV | Our calculation |
| Wan 2.2 A14B video at 720P | 59.8 GB peak on one GPU with offload. ComfyUI needs both 14.3 GB fp8 expert files | Yes | Published (Wan, Comfy-Org files) |
| FLUX.2 [dev], fp8 transformer plus fp8 text encoder | 35.5 GB + 18.0 GB = 53.5 GB of files | Yes, both in memory at once | Published file sizes, our sum |
| FLUX.1 [dev], NVFP4 | 9.2 GB file | Yes, on a card that can compute in FP4 | Published (Black Forest Labs) |
| LTX-2 audio and video, NVFP4 plus fp8 text encoder | 20.0 GB + 13.2 GB = 33.2 GB of files | Yes | Published file sizes, our sum |
| QLoRA fine-tuning of gpt-oss-120b | 65 GB | Yes | Published (Unsloth) |
| 16-bit LoRA, 30B to 34B | 64 to 80 GB | Yes | Published (Axolotl) |
FP4 is the other reason to pick this card. vLLM runs NVFP4 checkpoints natively on Blackwell, and ComfyUI computes in NVFP4 on compute capability 10 and newer when PyTorch is a cu130 build. Two licence notes before you build a product on these models: FLUX [dev]-family outputs can be used commercially, but running the model itself in a commercial service needs a licence from Black Forest Labs, and LTX-2 needs a paid licence from $10M of annual revenue. The gpt-oss GPU requirements guide and FLUX in ComfyUI go deeper on gpt-oss and FLUX.
Measured on QuantaCloud#
Across the catalog, most single-GPU VMs are running in about 3 minutes (median).
RTX PRO 6000 vs H100 PCIe vs H200 NVL#
The RTX PRO 6000 sits between the H100 PCIe and the H200 NVL on memory and below both on bandwidth.
| RTX PRO 6000 Blackwell | H100 PCIe | H200 NVL | |
|---|---|---|---|
| Memory | 96 GB GDDR7 | 80 GB HBM2e | 141 GB HBM3e |
| Bandwidth | 1,597 or 1,792 GB/s, by edition | 2,000 GB/s | 4.8 TB/s |
| FP8 and FP4 | Both | FP8 only | FP8 only |
| NVLink on the card | Not supported | 2-card bridge | 2- or 4-way bridge |
| Compute capability | 12.0 | 9.0 | 9.0 |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Choose the RTX PRO 6000 when the job needs 80 to 96 GB on one GPU, or when you want FP4, which neither Hopper card has. Choose the H100 PCIe when the model fits in 80 GB and you want Hopper's bandwidth and kernels (FlashAttention-3 targets H100-class GPUs). Choose the H200 NVL when you need more than 96 GB on one GPU, or the most bandwidth for serving 70B to 120B models. The RTX PRO 6000 vs H100, RTX PRO 6000 vs H200 and RTX PRO 6000 vs A100 comparisons go through each pair in detail.
Launch the RTX PRO 6000 with a template#
Pick the template by the first job you will run. The template does not change the price, and none of them comes with models preinstalled.
| Template | What you get on Blackwell | Launch |
|---|---|---|
| Bare Metal | Ubuntu 22.04, the NVIDIA driver and Docker over SSH. Pick your vLLM image tag from the driver table above | Ubuntu + Docker on RTX PRO 6000 |
| PyTorch + Jupyter | JupyterLab with PyTorch behind your QuantaCloud login. Run the architecture check in the first cell | Jupyter on RTX PRO 6000 |
| Open WebUI + Ollama | A private chat UI. Ollama's GPU docs list compute capability 12.0, and Open WebUI has its own sign-in | Open WebUI on RTX PRO 6000 |
| ComfyUI | Image and video workflows. Upload or download your models after launch | ComfyUI on RTX PRO 6000 |
The templates docs list what each template contains, and connecting over SSH covers the Bare Metal login.
FAQ#
Is it the Server Edition or the Workstation Edition?
The catalog does not say, but nvidia-smi -q on the VM does. The difference that matters is bandwidth: 1,597 GB/s on the Server Edition and 1,792 GB/s on the Workstation and Max-Q.
Does the RTX PRO 6000 have NVLink?
No. NVIDIA lists NVLink as not supported, so the GPUs in a 2x or 4x VM talk over PCIe. To split one model across them, vLLM's docs recommend pipeline parallelism over tensor parallelism. The InfiniBand vs NVLink guide explains what each link is for.
Can I use MIG on it?
Not as a catalog option: on 2026-09-27 every RTX PRO 6000 offer was a full 96 GB GPU. The card itself supports up to four 24 GB MIG instances.
Does FP4 work in vLLM and ComfyUI?
Yes, with current builds. vLLM runs NVFP4 checkpoints natively on Blackwell. ComfyUI needs a cu130 PyTorch build for its fast NVFP4 path, and without one its developers say NVFP4 can be up to 2x slower than fp8.
Why does my old container fail on it?
It was built for CUDA 12.6 or older, which has no Blackwell kernels. Rebuild it on a CUDA 12.8 or newer base with a cu128, cu129 or cu130 PyTorch that matches the driver (cu130 needs 580 or newer), then check that torch.cuda.get_arch_list() includes sm_120.
Is it a VM or bare metal?
A VM with Ubuntu 22.04, the NVIDIA driver and Docker. "Bare Metal" is the name of the plain Ubuntu template, not of the hardware.
What happens to my data when I stop?
It is deleted with the VM. Stopping terminates the instance and deletes its disk, and there are no volumes or snapshots, so copy checkpoints and outputs off first. The unused seconds of the current hour are refunded.
Can I get dedicated RTX PRO 6000 servers?
Yes, as reserved capacity built to order. We order and build the servers to your spec, and the configuration, lead time and terms are quoted in writing. Dedicated GPU servers covers the options, or you can send a capacity brief.
My rule for the RTX PRO 6000 is this: rent it when the job needs more than 48 GB and fits in 96 GB on one GPU, or when you want FP4, and check the driver before you install anything. If the model needs more than 96 GB on one GPU, the H200 NVL is the next step. Launch an RTX PRO 6000 , run nvidia-smi, and pick your PyTorch and vLLM builds from the driver table.