The NVIDIA H200 NVL is the biggest GPU in QuantaCloud's on-demand catalog: 141 GB of HBM3e at 4.8 TB/s on one card, from the console price. That is enough memory to serve a 70B model at FP8 with a full 128k-token context on a single GPU, or gpt-oss-120b with room for 15 long conversations at once. Every instance is a QuantaCloud VM with one or more H200 NVL GPUs, paid from prepaid credit, and the unused seconds of the current hour come back to your balance when you stop.
On 2026-09-27 the catalog listed 1x and 2x configurations in us-east-1 (Virginia), and a 4x configuration appeared 13 minutes later. The table above is the live list.
H200 NVL vs H200 SXM specs#
The memory is identical: both versions of the H200 carry 141 GB of HBM3e at 4.8 TB/s. What differs is the package, the power limit and the GPU-to-GPU link, and the SXM's higher power buys it more peak throughput.
| Spec | H200 NVL (on demand) | H200 SXM (built to order) |
|---|---|---|
| Architecture | Hopper, compute capability 9.0 | Hopper, compute capability 9.0 |
| GPU memory | 141 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 4.8 TB/s | 4.8 TB/s |
| FP8 Tensor Core | 3,341 TFLOPS | 3,958 TFLOPS |
| BF16 and FP16 Tensor Core | 1,671 TFLOPS | 1,979 TFLOPS |
| TF32 Tensor Core | 835 TFLOPS | 989 TFLOPS |
| FP64 | 30 TFLOPS | 34 TFLOPS |
| FP4 | No | No |
| Max power | Up to 600 W, configurable | Up to 700 W, configurable |
| Form factor | PCIe card, dual-slot, air-cooled | SXM module |
| GPU to GPU | 2- or 4-way NVLink bridge, 900 GB/s per GPU | NVLink, 900 GB/s per GPU |
| Host link | PCIe Gen5, 128 GB/s | PCIe Gen5, 128 GB/s |
| Multi-Instance GPU | Up to 7 instances | Up to 7 instances |
| Decoders | 7 NVDEC, 7 JPEG | 7 NVDEC, 7 JPEG |
| NVIDIA server platform | MGX H200 NVL, up to 8 GPUs | HGX H200, 4 or 8 GPUs |
NVIDIA quotes Tensor Core figures with sparsity, so dense throughput is half of each Tensor Core number in the table.
For single-GPU inference the gap is small: generating tokens at small batch sizes is bound by memory bandwidth, and that is 4.8 TB/s on both. The SXM's extra tensor throughput, about 18 percent (our calculation: 3,958 / 3,341 = 1.18), shows up in prefill and training. Its real advantage is the HGX baseboard, where NVSwitch connects all eight GPUs, while an NVL bridge links at most four cards. That makes eight-GPU training an HGX H200 job, and HGX H200 servers are reserved capacity that we build to order: send a capacity brief.
What fits in 141 GB#
The working budget is 129.2 GiB, not 141 GB: vLLM claims 92 percent of GPU memory by default through --gpu-memory-utilization, the card reports 143,771 MiB, and weights, KV cache and activations all have to fit inside the budget (our calculation: 143,771 / 1,024 x 0.92).
| Workload | Memory it needs | On one H200 NVL | Basis |
|---|---|---|---|
| gpt-oss-120b (MXFP4) | 65.2 GB (60.8 GiB) of weights, plus 4.5 GiB of KV cache per full 131,072-token sequence | Yes, with room for about 15 full-length sequences: (129.2 - 60.8) / 4.5 = 15.2 | Published sizes, our calculation |
| Llama-3.3-70B at FP8, one full 128k context | 72.7 GB (67.7 GiB) FP8 checkpoint + 40 GiB of BF16 KV cache = 107.7 GiB, or 115.6 GB | Yes | Published size, our calculation |
| Llama-3.3-70B at BF16 | 141.1 GB (131.4 GiB) of weights alone | No. Use FP8, or 2x H200 NVL | Our calculation |
| Qwen3-32B at BF16, 32k context | 65.5 GB (61.0 GiB) of weights, plus 8 GiB of KV cache per 32,768-token sequence | Yes, about 8 sequences: (129.2 - 61.0) / 8 = 8.5 | Our calculation |
| QLoRA fine-tune, 70B model | 41 GB (Unsloth), 40 to 48 GB (Axolotl) | Yes, with room for longer sequences | Published |
| QLoRA fine-tune, gpt-oss-120b | 65 GB (Unsloth) | Yes | Published |
| 16-bit LoRA, 32B model | 76 GB (Unsloth), 64 to 80 GB (Axolotl) | Yes | Published |
| 16-bit LoRA, 70B model | 164 GB (Unsloth) | No. It needs 2x H200 NVL | Published |
The KV figures come from each model's config: 320 KiB per token for Llama-3.3-70B, 256 KiB for Qwen3-32B, and 36 KiB for gpt-oss-120b, whose sliding-window layers stop growing after 128 tokens. The sequence counts leave out activations and CUDA graph memory, so treat them as ceilings. --kv-cache-dtype fp8 halves every KV number, and the KV cache guide works through the formula. For model-by-model detail, see gpt-oss GPU requirements, DeepSeek GPU requirements and the LoRA, QLoRA and full fine-tuning VRAM guide.
Two H200 NVL in one VM#
Two cards give you 282 GB, enough for the jobs that miss on one. Llama-3.3-70B at BF16 fits with about 127 GiB left for KV cache, three full 128k-token sequences (our calculation: 2 x 129.17 - 131.42 = 126.9, and 126.9 / 40 = 3.2), and 16-bit LoRA on a 70B model, which Unsloth puts at 164 GB, fits once FSDP or DeepSpeed shards the model across both cards. In vLLM you split a model across both cards with --tensor-parallel-size 2.
The open question is the link between the cards. The H200 NVL supports an NVLink bridge at 900 GB/s per GPU, and QuantaCloud's API flags these offers as NVLink, but I do not treat that as settled until nvidia-smi topo -m on the VM shows NV links between the two GPUs. When a VM shows PCIe only, vLLM's own guidance is pipeline parallelism (--pipeline-parallel-size 2) rather than tensor parallelism.
H200 NVL vs H100 PCIe, RTX PRO 6000 and A100#
The H200 NVL is the right pick when the model and its context need more than 96 GB on one GPU, or when tokens per second per user matter more than the hourly price. It has 76 percent more memory than any 80 GB GPU and 2.4 times the bandwidth of the H100 PCIe (our calculation: 141 / 80 = 1.76 and 4.8 / 2.0 = 2.4). It also runs the same Hopper code as the H100, since both are compute capability 9.0. The RTX PRO 6000 vs H200 and A100 vs H200 comparisons go through two of these pairs in detail.
| GPU | Memory | Bandwidth | FP8 / FP4 | Choose it when |
|---|---|---|---|---|
| H200 NVL | 141 GB HBM3e | 4.8 TB/s | Yes / No | The job needs more than 96 GB on one GPU, or the fastest decoding |
| RTX PRO 6000 Blackwell | 96 GB GDDR7 | 1,597 or 1,792 GB/s, by edition | Yes / Yes | The job fits in 96 GB, or you serve NVFP4 checkpoints |
| H100 PCIe | 80 GB HBM2e | 2.0 TB/s | Yes / No | The job fits in 80 GB with its context and needs FP8 |
| A100 SXM4 80GB | 80 GB HBM2e | 2,039 GB/s | No / No | BF16 inference or fine-tuning that fits in 80 GB |
| GPU | Memory | From | Available now |
|---|---|---|---|
| H200 NVL | - | Not listed | No |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
The comparison I run is the cost of holding the model, not the price of a GPU-hour. On 2026-09-27, Llama-3.3-70B at FP8 with a full 128k context needed one H200 NVL at $3.43 an hour or a 2x H100 PCIe VM at $5.18, so the H200 NVL held it for 34 percent less per hour (our calculation: 1 - 3.43 / 5.18 = 0.34). A 2x A100 PCIe VM held it for less, $2.95 an hour, but the A100 has no FP8 Tensor Cores, so vLLM runs the FP8 checkpoint weight-only there. When the model fits on one 80 GB card with room to spare, check the H100 PCIe and A100 rows in the live table before you pay for 141 GB. Past 141 GB per GPU, the next step is Blackwell: the B200 has 180 GB and is built to order. H200 vs B200 compares the two for training and inference. If you are pricing a purchase instead, the H200 price guide works through buying against renting.
Measured on QuantaCloud#
Across QuantaCloud, most single-GPU VMs are running in about 3 minutes (median). Check the NVIDIA driver version before you pull any image: vLLM's default image is a CUDA 13 build that needs an R580 or newer driver, and on anything older you use its -cu129 tag instead.
Launch it with a template#
The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.
| Template | What I would run on an H200 NVL | Launch |
|---|---|---|
| Bare Metal: Ubuntu 22.04, NVIDIA driver, Docker | vLLM in Docker serving gpt-oss-120b or a 70B model at FP8, following the vLLM Docker guide | Launch Bare Metal |
| PyTorch + Jupyter | QLoRA on a 70B model, or 16-bit LoRA up to 32B, from a Jupyter notebook | Launch PyTorch + Jupyter |
| Open WebUI + Ollama | A private chat on gpt-oss-120b, which leaves most of the card free for context | Launch Open WebUI |
| ComfyUI | Wan 2.2 A14B video at 720p, which peaked at 59.8 GB in Wan's own single-GPU test | Launch ComfyUI |
Every instance is a VM: Bare Metal is the name of the plain Ubuntu template, not bare-metal hardware. SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login, and Open WebUI also asks you to create its own admin account on the first visit. Download the model weights you need after boot (fast Hugging Face downloads), and remember they are deleted with the disk when you stop. The templates docs list what each one contains.
FAQ#
Is the H200 NVL the same GPU as the H200 SXM?
Same memory, different package. Both have 141 GB of HBM3e at 4.8 TB/s. The NVL is a PCIe card limited to 600 W that bridges to one or three other cards, and the SXM runs at up to 700 W in HGX servers. QuantaCloud offers the NVL on demand and builds HGX H200 servers to order.
How much of the 141 GB can I use?
Plan on less than 141 GB. By default, vLLM claims 92 percent of the GPU memory that nvidia-smi reports. I size a model plus its KV cache against about 129 GiB and leave the rest for the CUDA context.
Can I get four or eight H200 NVL in one VM?
One and two are the usual configurations, and a 4x configuration appeared in the catalog on 2026-09-27, so check the live table. There is no 8x H200 on demand. Eight H200 GPUs in one server is an HGX H200 or MGX H200 NVL build, which is reserved capacity.
Is it available right now?
The configuration table at the top is live, and availability changes by the minute. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 the H200 NVL was in us-east-1 (Virginia).
What happens when I stop the instance?
Stopping terminates the VM and deletes its disk. There are no volumes or snapshots, so copy checkpoints and outputs off first. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. Pricing and the billing docs have the full rules.
Can I rent HGX H200 servers or an H200 cluster?
Yes, as reserved capacity built to order. Tell us the GPU count, network and term, and we return the configuration, lead time and terms in writing. Send a capacity brief, or see how multi-node builds work on the GPU clusters page.
The rule I follow for the H200 NVL: rent it when the model and its KV cache need more than 96 GB on one GPU, or when you would otherwise split a model across two 80 GB cards. Below that line, check the H100 PCIe and the RTX PRO 6000 in the GPU catalog first. If one H200 NVL fits your job, start with a 1x VM, run the vLLM Docker guide against your own prompts, and check tokens per second before you add a second card. Launch H200 NVL