On-demand GPU

NVIDIA H200 NVL

141 GB. Room to build.

H200 NVL cloud GPUs with 141 GB HBM3e and 4.8 TB/s: 1 or 2 per VM, live hourly prices, and what fits in 141 GB, from 70B at FP8 to gpt-oss-120b.

Full SSH accessYour choice of template
Starting fromSee console
Explore configurations
Check availabilitySee configurations
NVIDIA H200 NVLHOPPER
VRAMVRAMVRAMVRAMVRAMVRAMHOPPER
Room for the work ahead.141 GB HBM3e
GPU memory
141 GB HBM3e
Memory bandwidth
4.8 TB/s
Architecture
Hopper
Deployment
On demand

Your GPU. Your configuration.

Choose where you start.

0 configurations

See current H200 NVL configurations in the console.

Prices checked

The NVIDIA H200 NVL is the biggest GPU in QuantaCloud's on-demand catalog: 141 GB of HBM3e at 4.8 TB/s on one card, from the console price. That is enough memory to serve a 70B model at FP8 with a full 128k-token context on a single GPU, or gpt-oss-120b with room for 15 long conversations at once. Every instance is a QuantaCloud VM with one or more H200 NVL GPUs, paid from prepaid credit, and the unused seconds of the current hour come back to your balance when you stop.

On 2026-09-27 the catalog listed 1x and 2x configurations in us-east-1 (Virginia), and a 4x configuration appeared 13 minutes later. The table above is the live list.

H200 NVL vs H200 SXM specs#

The memory is identical: both versions of the H200 carry 141 GB of HBM3e at 4.8 TB/s. What differs is the package, the power limit and the GPU-to-GPU link, and the SXM's higher power buys it more peak throughput.

SpecH200 NVL (on demand)H200 SXM (built to order)
ArchitectureHopper, compute capability 9.0Hopper, compute capability 9.0
GPU memory141 GB HBM3e141 GB HBM3e
Memory bandwidth4.8 TB/s4.8 TB/s
FP8 Tensor Core3,341 TFLOPS3,958 TFLOPS
BF16 and FP16 Tensor Core1,671 TFLOPS1,979 TFLOPS
TF32 Tensor Core835 TFLOPS989 TFLOPS
FP6430 TFLOPS34 TFLOPS
FP4NoNo
Max powerUp to 600 W, configurableUp to 700 W, configurable
Form factorPCIe card, dual-slot, air-cooledSXM module
GPU to GPU2- or 4-way NVLink bridge, 900 GB/s per GPUNVLink, 900 GB/s per GPU
Host linkPCIe Gen5, 128 GB/sPCIe Gen5, 128 GB/s
Multi-Instance GPUUp to 7 instancesUp to 7 instances
Decoders7 NVDEC, 7 JPEG7 NVDEC, 7 JPEG
NVIDIA server platformMGX H200 NVL, up to 8 GPUsHGX H200, 4 or 8 GPUs

NVIDIA quotes Tensor Core figures with sparsity, so dense throughput is half of each Tensor Core number in the table.

For single-GPU inference the gap is small: generating tokens at small batch sizes is bound by memory bandwidth, and that is 4.8 TB/s on both. The SXM's extra tensor throughput, about 18 percent (our calculation: 3,958 / 3,341 = 1.18), shows up in prefill and training. Its real advantage is the HGX baseboard, where NVSwitch connects all eight GPUs, while an NVL bridge links at most four cards. That makes eight-GPU training an HGX H200 job, and HGX H200 servers are reserved capacity that we build to order: send a capacity brief.

What fits in 141 GB#

The working budget is 129.2 GiB, not 141 GB: vLLM claims 92 percent of GPU memory by default through --gpu-memory-utilization, the card reports 143,771 MiB, and weights, KV cache and activations all have to fit inside the budget (our calculation: 143,771 / 1,024 x 0.92).

WorkloadMemory it needsOn one H200 NVLBasis
gpt-oss-120b (MXFP4)65.2 GB (60.8 GiB) of weights, plus 4.5 GiB of KV cache per full 131,072-token sequenceYes, with room for about 15 full-length sequences: (129.2 - 60.8) / 4.5 = 15.2Published sizes, our calculation
Llama-3.3-70B at FP8, one full 128k context72.7 GB (67.7 GiB) FP8 checkpoint + 40 GiB of BF16 KV cache = 107.7 GiB, or 115.6 GBYesPublished size, our calculation
Llama-3.3-70B at BF16141.1 GB (131.4 GiB) of weights aloneNo. Use FP8, or 2x H200 NVLOur calculation
Qwen3-32B at BF16, 32k context65.5 GB (61.0 GiB) of weights, plus 8 GiB of KV cache per 32,768-token sequenceYes, about 8 sequences: (129.2 - 61.0) / 8 = 8.5Our calculation
QLoRA fine-tune, 70B model41 GB (Unsloth), 40 to 48 GB (Axolotl)Yes, with room for longer sequencesPublished
QLoRA fine-tune, gpt-oss-120b65 GB (Unsloth)YesPublished
16-bit LoRA, 32B model76 GB (Unsloth), 64 to 80 GB (Axolotl)YesPublished
16-bit LoRA, 70B model164 GB (Unsloth)No. It needs 2x H200 NVLPublished

The KV figures come from each model's config: 320 KiB per token for Llama-3.3-70B, 256 KiB for Qwen3-32B, and 36 KiB for gpt-oss-120b, whose sliding-window layers stop growing after 128 tokens. The sequence counts leave out activations and CUDA graph memory, so treat them as ceilings. --kv-cache-dtype fp8 halves every KV number, and the KV cache guide works through the formula. For model-by-model detail, see gpt-oss GPU requirements, DeepSeek GPU requirements and the LoRA, QLoRA and full fine-tuning VRAM guide.

Two H200 NVL in one VM#

Two cards give you 282 GB, enough for the jobs that miss on one. Llama-3.3-70B at BF16 fits with about 127 GiB left for KV cache, three full 128k-token sequences (our calculation: 2 x 129.17 - 131.42 = 126.9, and 126.9 / 40 = 3.2), and 16-bit LoRA on a 70B model, which Unsloth puts at 164 GB, fits once FSDP or DeepSpeed shards the model across both cards. In vLLM you split a model across both cards with --tensor-parallel-size 2.

The open question is the link between the cards. The H200 NVL supports an NVLink bridge at 900 GB/s per GPU, and QuantaCloud's API flags these offers as NVLink, but I do not treat that as settled until nvidia-smi topo -m on the VM shows NV links between the two GPUs. When a VM shows PCIe only, vLLM's own guidance is pipeline parallelism (--pipeline-parallel-size 2) rather than tensor parallelism.

H200 NVL vs H100 PCIe, RTX PRO 6000 and A100#

The H200 NVL is the right pick when the model and its context need more than 96 GB on one GPU, or when tokens per second per user matter more than the hourly price. It has 76 percent more memory than any 80 GB GPU and 2.4 times the bandwidth of the H100 PCIe (our calculation: 141 / 80 = 1.76 and 4.8 / 2.0 = 2.4). It also runs the same Hopper code as the H100, since both are compute capability 9.0. The RTX PRO 6000 vs H200 and A100 vs H200 comparisons go through two of these pairs in detail.

GPUMemoryBandwidthFP8 / FP4Choose it when
H200 NVL141 GB HBM3e4.8 TB/sYes / NoThe job needs more than 96 GB on one GPU, or the fastest decoding
RTX PRO 6000 Blackwell96 GB GDDR71,597 or 1,792 GB/s, by editionYes / YesThe job fits in 96 GB, or you serve NVFP4 checkpoints
H100 PCIe80 GB HBM2e2.0 TB/sYes / NoThe job fits in 80 GB with its context and needs FP8
A100 SXM4 80GB80 GB HBM2e2,039 GB/sNo / NoBF16 inference or fine-tuning that fits in 80 GB
GPUMemoryFromAvailable now
H200 NVL-Not listedNo
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes

The comparison I run is the cost of holding the model, not the price of a GPU-hour. On 2026-09-27, Llama-3.3-70B at FP8 with a full 128k context needed one H200 NVL at $3.43 an hour or a 2x H100 PCIe VM at $5.18, so the H200 NVL held it for 34 percent less per hour (our calculation: 1 - 3.43 / 5.18 = 0.34). A 2x A100 PCIe VM held it for less, $2.95 an hour, but the A100 has no FP8 Tensor Cores, so vLLM runs the FP8 checkpoint weight-only there. When the model fits on one 80 GB card with room to spare, check the H100 PCIe and A100 rows in the live table before you pay for 141 GB. Past 141 GB per GPU, the next step is Blackwell: the B200 has 180 GB and is built to order. H200 vs B200 compares the two for training and inference. If you are pricing a purchase instead, the H200 price guide works through buying against renting.

Measured on QuantaCloud#

Across QuantaCloud, most single-GPU VMs are running in about 3 minutes (median). Check the NVIDIA driver version before you pull any image: vLLM's default image is a CUDA 13 build that needs an R580 or newer driver, and on anything older you use its -cu129 tag instead.

Launch it with a template#

The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.

TemplateWhat I would run on an H200 NVLLaunch
Bare Metal: Ubuntu 22.04, NVIDIA driver, DockervLLM in Docker serving gpt-oss-120b or a 70B model at FP8, following the vLLM Docker guideLaunch Bare Metal
PyTorch + JupyterQLoRA on a 70B model, or 16-bit LoRA up to 32B, from a Jupyter notebookLaunch PyTorch + Jupyter
Open WebUI + OllamaA private chat on gpt-oss-120b, which leaves most of the card free for contextLaunch Open WebUI
ComfyUIWan 2.2 A14B video at 720p, which peaked at 59.8 GB in Wan's own single-GPU testLaunch ComfyUI

Every instance is a VM: Bare Metal is the name of the plain Ubuntu template, not bare-metal hardware. SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login, and Open WebUI also asks you to create its own admin account on the first visit. Download the model weights you need after boot (fast Hugging Face downloads), and remember they are deleted with the disk when you stop. The templates docs list what each one contains.

FAQ#

Is the H200 NVL the same GPU as the H200 SXM?

Same memory, different package. Both have 141 GB of HBM3e at 4.8 TB/s. The NVL is a PCIe card limited to 600 W that bridges to one or three other cards, and the SXM runs at up to 700 W in HGX servers. QuantaCloud offers the NVL on demand and builds HGX H200 servers to order.

How much of the 141 GB can I use?

Plan on less than 141 GB. By default, vLLM claims 92 percent of the GPU memory that nvidia-smi reports. I size a model plus its KV cache against about 129 GiB and leave the rest for the CUDA context.

Can I get four or eight H200 NVL in one VM?

One and two are the usual configurations, and a 4x configuration appeared in the catalog on 2026-09-27, so check the live table. There is no 8x H200 on demand. Eight H200 GPUs in one server is an HGX H200 or MGX H200 NVL build, which is reserved capacity.

Is it available right now?

The configuration table at the top is live, and availability changes by the minute. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 the H200 NVL was in us-east-1 (Virginia).

What happens when I stop the instance?

Stopping terminates the VM and deletes its disk. There are no volumes or snapshots, so copy checkpoints and outputs off first. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. Pricing and the billing docs have the full rules.

Can I rent HGX H200 servers or an H200 cluster?

Yes, as reserved capacity built to order. Tell us the GPU count, network and term, and we return the configuration, lead time and terms in writing. Send a capacity brief, or see how multi-node builds work on the GPU clusters page.


The rule I follow for the H200 NVL: rent it when the model and its KV cache need more than 96 GB on one GPU, or when you would otherwise split a model across two 80 GB cards. Below that line, check the H100 PCIe and the RTX PRO 6000 in the GPU catalog first. If one H200 NVL fits your job, start with a 1x VM, run the vLLM Docker guide against your own prompts, and check tokens per second before you add a second card. Launch H200 NVL

Keep building

Choose your next step.