On-demand GPU

NVIDIA RTX PRO 6000 Blackwell

96 GB. Room to build.

Rent RTX PRO 6000 Blackwell GPUs (96 GB GDDR7, FP4): 1, 2 or 4 per VM, live hourly prices, and the CUDA and PyTorch builds that run on Blackwell.

Full SSH accessYour choice of template
Starting from$2.39/ GPU-hr
Launch RTX PRO 6000 Blackwell
Available nowSee configurations
NVIDIA RTX PRO 6000 BlackwellBLACKWELL
VRAMVRAMVRAMVRAMVRAMVRAMBLACKWELL
Room for the work ahead.96 GB GDDR7
GPU memory
96 GB GDDR7
Memory bandwidth
1,597–1,792 GB/s by edition
Architecture
Blackwell
Deployment
On demand

Your GPU. Your configuration.

Choose where you start.

2 configurations
1 GPULowest price

1x RTX PRO 6000 Blackwell

Virginia · us-east-1

vCPU
16
RAM
144 GB
Disk
725 GB
$2.39/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX PRO 6000 Blackwell

us-midwest-4

vCPU
16
RAM
144 GB
Disk
725 GB
$2.39/ instance-hour
Available nowLaunch configuration

Prices checked

The RTX PRO 6000 Blackwell is the GPU to rent when a job needs 96 GB on one card. On QuantaCloud, RTX PRO 6000 cloud VMs run from $2.39/GPU-hr, with 1, 2 or 4 cards per VM, and on 2026-09-27 it was the only GPU in the on-demand catalog with FP4 Tensor Cores. The catch is software: Blackwell needs CUDA 12.8 or newer builds, and many older containers fail on it.

Each configuration is a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu, and its local disk is part of the hourly price. On 2026-09-27 the 1x came with 16 vCPUs, 144 GB of RAM and 725 GB of disk, and the 4x with 60 vCPUs, 576 GB of RAM and 2,900 GB, in us-east-1 (Virginia) and us-midwest-4. The first hour is charged at launch and the unused seconds of the current hour are refunded when you stop. The pricing page has the rules, and the GPU catalog has every other card.

Which RTX PRO 6000 this is#

The edition matters more on this GPU than on any other in the catalog. NVIDIA sells three RTX PRO 6000 Blackwell editions with the same 96 GB and different power and bandwidth:

EditionCoolingPowerMemory bandwidth
Server EditionPassive, for server airflowUp to 600 W, configurable (600 W or 450 W mode, 300 W minimum)1,597 GB/s
Workstation EditionFlow-through fan600 W1,792 GB/s
Max-Q Workstation EditionBlower fan300 W1,792 GB/s

The catalog lists the card only as "RTX PRO 6000 Blackwell", and nvidia-smi -q on any VM you launch shows which edition it is.

RTX PRO 6000 Blackwell specs#

The RTX PRO 6000 Blackwell is a 96 GB card with FP4 Tensor Cores and no NVLink. The figures below are from NVIDIA's Server Edition product page and datasheet, with the dense Tensor rates worked out from NVIDIA's architecture whitepaper. All three editions share the 96 GB, compute capability 12.0 and the four-instance MIG limit.

SpecRTX PRO 6000 Blackwell Server Edition
ArchitectureNVIDIA Blackwell (GB202), compute capability 12.0
GPU memory96 GB GDDR7 with ECC, 512-bit
Memory bandwidth1,597 GB/s (Workstation and Max-Q: 1,792 GB/s)
CUDA cores24,064
Tensor Cores752, fifth generation, with FP8 and FP4
RT Cores188, fourth generation
FP32120 TFLOPS
FP8 and FP4 Tensor2 PFLOPS and 4 PFLOPS with sparsity. Dense, about 936 and 1,871 TFLOPS at the Server Edition's clock (our calculation), and 1,007.6 and 2,015.2 on the Workstation Edition
NVLinkNot supported
MIGUp to 4 instances of 24 GB
System interfacePCIe 5.0 x16
PowerUp to 600 W, configurable
Video engines4 NVENC, 4 NVDEC, 4 JPEG

Do not confuse it with the RTX 6000 Ada (Ada, compute capability 8.9) or the RTX A6000 (Ampere, 8.6). Both of those are 48 GB cards, half the memory of the PRO 6000 Blackwell. For the RTX 5090, see RTX PRO 6000 vs RTX 5090.

Software that runs on Blackwell#

The core problem with Blackwell is software, not hardware. It needs CUDA 12.8 or newer builds, and PyTorch builds for CUDA 12.6 or older have no kernels for it. When a stack is too old, you see one of these errors:

Output
... with CUDA capability sm_120 is not compatible with the current PyTorch installation.
... no kernel image is available for execution on the device

Two commands tell you where you stand. The first prints the driver version, and the second must list sm_120:

Terminal
nvidia-smi
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"

The driver decides which builds run. This is the table I work from:

Driver on the VMPyTorchvLLM 0.30.0 Docker imageComfyUI v0.37
580 or newerpip install torch works: on Linux it installs CUDA 13.0 wheelsvllm/vllm-openai:v0.30.0Runs. It requires a cu130 PyTorch build
570 to 575Plain pip install torch installs CUDA 13.0 wheels this driver cannot run. Pin a cu129 build (Linux, PyTorch 2.13 and older) or cu128 (2.11 and older)vllm/vllm-openai:v0.30.0-cu129Needs a 580 driver first
Older than 570No Blackwell supportNot supportedNot supported

A cu129 pin looks like pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu129. The driver and CUDA version guide explains the checks, and installing vLLM and the vLLM Docker guide show where the tag goes.

What 96 GB lets you run#

The rule I follow is the same as on any card: weights plus KV cache under 92% of memory, which is about 88 GiB here (our calculation: the card reports 97,887 MiB, and 97,887 / 1,024 x 0.92 = 87.9 GiB, vLLM's default).

WorkloadMemory it needsOn one RTX PRO 6000Basis
gpt-oss-120bAbout 65 GB (60.8 GiB) on disk, plus 4.5 GiB of KV per full 131k-token sequenceYes, with KV room for about six full-length sequences at oncePublished sizes, our calculation
Qwen3-32B in BF16, one 32k sequenceAbout 69 GiBYes, with about 19 GiB to spareOur calculation
Llama-3.3-70B in FP872.7 GB (67.7 GiB) checkpoint, leaving about 20.3 GiB for KVTight: about 66k tokens of BF16 KV, or twice that with FP8 KVOur calculation
Wan 2.2 A14B video at 720P59.8 GB peak on one GPU with offload. ComfyUI needs both 14.3 GB fp8 expert filesYesPublished (Wan, Comfy-Org files)
FLUX.2 [dev], fp8 transformer plus fp8 text encoder35.5 GB + 18.0 GB = 53.5 GB of filesYes, both in memory at oncePublished file sizes, our sum
FLUX.1 [dev], NVFP49.2 GB fileYes, on a card that can compute in FP4Published (Black Forest Labs)
LTX-2 audio and video, NVFP4 plus fp8 text encoder20.0 GB + 13.2 GB = 33.2 GB of filesYesPublished file sizes, our sum
QLoRA fine-tuning of gpt-oss-120b65 GBYesPublished (Unsloth)
16-bit LoRA, 30B to 34B64 to 80 GBYesPublished (Axolotl)

FP4 is the other reason to pick this card. vLLM runs NVFP4 checkpoints natively on Blackwell, and ComfyUI computes in NVFP4 on compute capability 10 and newer when PyTorch is a cu130 build. Two licence notes before you build a product on these models: FLUX [dev]-family outputs can be used commercially, but running the model itself in a commercial service needs a licence from Black Forest Labs, and LTX-2 needs a paid licence from $10M of annual revenue. The gpt-oss GPU requirements guide and FLUX in ComfyUI go deeper on gpt-oss and FLUX.

Measured on QuantaCloud#

Across the catalog, most single-GPU VMs are running in about 3 minutes (median).

RTX PRO 6000 vs H100 PCIe vs H200 NVL#

The RTX PRO 6000 sits between the H100 PCIe and the H200 NVL on memory and below both on bandwidth.

RTX PRO 6000 BlackwellH100 PCIeH200 NVL
Memory96 GB GDDR780 GB HBM2e141 GB HBM3e
Bandwidth1,597 or 1,792 GB/s, by edition2,000 GB/s4.8 TB/s
FP8 and FP4BothFP8 onlyFP8 only
NVLink on the cardNot supported2-card bridge2- or 4-way bridge
Compute capability12.09.09.0
GPUMemoryFromAvailable now
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
H200 NVL-Not listedNo

Choose the RTX PRO 6000 when the job needs 80 to 96 GB on one GPU, or when you want FP4, which neither Hopper card has. Choose the H100 PCIe when the model fits in 80 GB and you want Hopper's bandwidth and kernels (FlashAttention-3 targets H100-class GPUs). Choose the H200 NVL when you need more than 96 GB on one GPU, or the most bandwidth for serving 70B to 120B models. The RTX PRO 6000 vs H100, RTX PRO 6000 vs H200 and RTX PRO 6000 vs A100 comparisons go through each pair in detail.

Launch the RTX PRO 6000 with a template#

Pick the template by the first job you will run. The template does not change the price, and none of them comes with models preinstalled.

TemplateWhat you get on BlackwellLaunch
Bare MetalUbuntu 22.04, the NVIDIA driver and Docker over SSH. Pick your vLLM image tag from the driver table aboveUbuntu + Docker on RTX PRO 6000
PyTorch + JupyterJupyterLab with PyTorch behind your QuantaCloud login. Run the architecture check in the first cellJupyter on RTX PRO 6000
Open WebUI + OllamaA private chat UI. Ollama's GPU docs list compute capability 12.0, and Open WebUI has its own sign-inOpen WebUI on RTX PRO 6000
ComfyUIImage and video workflows. Upload or download your models after launchComfyUI on RTX PRO 6000

The templates docs list what each template contains, and connecting over SSH covers the Bare Metal login.

FAQ#

Is it the Server Edition or the Workstation Edition?

The catalog does not say, but nvidia-smi -q on the VM does. The difference that matters is bandwidth: 1,597 GB/s on the Server Edition and 1,792 GB/s on the Workstation and Max-Q.

No. NVIDIA lists NVLink as not supported, so the GPUs in a 2x or 4x VM talk over PCIe. To split one model across them, vLLM's docs recommend pipeline parallelism over tensor parallelism. The InfiniBand vs NVLink guide explains what each link is for.

Can I use MIG on it?

Not as a catalog option: on 2026-09-27 every RTX PRO 6000 offer was a full 96 GB GPU. The card itself supports up to four 24 GB MIG instances.

Does FP4 work in vLLM and ComfyUI?

Yes, with current builds. vLLM runs NVFP4 checkpoints natively on Blackwell. ComfyUI needs a cu130 PyTorch build for its fast NVFP4 path, and without one its developers say NVFP4 can be up to 2x slower than fp8.

Why does my old container fail on it?

It was built for CUDA 12.6 or older, which has no Blackwell kernels. Rebuild it on a CUDA 12.8 or newer base with a cu128, cu129 or cu130 PyTorch that matches the driver (cu130 needs 580 or newer), then check that torch.cuda.get_arch_list() includes sm_120.

Is it a VM or bare metal?

A VM with Ubuntu 22.04, the NVIDIA driver and Docker. "Bare Metal" is the name of the plain Ubuntu template, not of the hardware.

What happens to my data when I stop?

It is deleted with the VM. Stopping terminates the instance and deletes its disk, and there are no volumes or snapshots, so copy checkpoints and outputs off first. The unused seconds of the current hour are refunded.

Can I get dedicated RTX PRO 6000 servers?

Yes, as reserved capacity built to order. We order and build the servers to your spec, and the configuration, lead time and terms are quoted in writing. Dedicated GPU servers covers the options, or you can send a capacity brief.


My rule for the RTX PRO 6000 is this: rent it when the job needs more than 48 GB and fits in 96 GB on one GPU, or when you want FP4, and check the driver before you install anything. If the model needs more than 96 GB on one GPU, the H200 NVL is the next step. Launch an RTX PRO 6000 , run nvidia-smi, and pick your PyTorch and vLLM builds from the driver table.

Keep building

Choose your next step.