The upgrade pays when the model outgrows 80 GB or the GPU serves FP8 to many users, and it does not when a BF16 job that fits in 80 GB spends its hours waiting on people. The H200 NVL has 141 GB of HBM3e at 4.8 TB/s and FP8 Tensor Cores. The A100 has 80 GB of HBM2e at up to 2,039 GB/s and no FP8. On 2026-09-27 a 1x H200 NVL cost $3.43 an hour and a 1x A100 SXM4 $1.50, 2.29 times as much, which is close to the ratio of their bandwidths, 2.35. So for plain BF16 generation on a model that fits both, they cost about the same per token, and the other differences decide the choice.
A100 and H200 NVL specs side by side#
The H200 NVL leads on every line except price and MIG, where both split into up to seven instances. The table carries both 80 GB A100s in the catalog, because the PCIe card is how QuantaCloud rents two or four of them.
| Spec | A100 80GB SXM4 | A100 80GB PCIe | H200 NVL |
|---|---|---|---|
| Architecture (compute capability) | Ampere (8.0) | Ampere (8.0) | Hopper (9.0) |
| GPU memory | 80 GB HBM2e | 80 GB HBM2e | 141 GB HBM3e |
| Memory bandwidth | 2,039 GB/s | 1,935 GB/s | 4.8 TB/s |
| FP32 | 19.5 TFLOPS | 19.5 TFLOPS | 60 TFLOPS |
| TF32 Tensor Core (dense) | 156 TFLOPS | 156 TFLOPS | 417.5 TFLOPS |
| BF16 and FP16 Tensor Core (dense) | 312 TFLOPS | 312 TFLOPS | 835.5 TFLOPS |
| FP8 Tensor Core (dense) | Not supported | Not supported | 1,670.5 TFLOPS |
| FP64 | 9.7 TFLOPS, 19.5 on Tensor Cores | 9.7 TFLOPS, 19.5 on Tensor Cores | 30 TFLOPS, 60 on Tensor Cores |
| GPU to GPU | NVLink, 600 GB/s through the HGX A100 board | NVLink bridge for 2 GPUs, 600 GB/s | 2- or 4-way NVLink bridge, 900 GB/s per GPU |
| Host link | PCIe Gen4, 64 GB/s | PCIe Gen4, 64 GB/s | PCIe Gen5, 128 GB/s |
| Multi-Instance GPU | Up to 7 x 10 GB | Up to 7 x 10 GB | Up to 7 |
| Max power | 400 W | 300 W | Up to 600 W, configurable |
| Form factor | SXM4 module | Dual-slot PCIe card | Dual-slot PCIe card, air-cooled |
The A100 figures come from NVIDIA's A100 datasheet and the H200 NVL figures from NVIDIA's H200 product page, which quotes Tensor Core throughput with sparsity, so the table halves those. On paper the H200 NVL reads memory 2.35 times as fast as the A100 SXM4 (our calculation: 4,800 / 2,039), and has 2.68 times the dense BF16 rate (835.5 / 312), 3.1 times the FP64 rate (30 / 9.7) and 1.76 times the memory (141 / 80). For FP8 checkpoints the gap is 5.4 times, because the A100 loads them weight-only and multiplies at its BF16 rate (1,670.5 / 312).
Prices, and why BF16 is close to a tie#
The A100 hour cost less than half as much on 2026-09-27, almost exactly in line with its bandwidth.
| GPU | Memory | From | Available now |
|---|---|---|---|
| A100 SXM4 80GB | 80 GB | $1.50/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 15:42 UTC
A 1x A100 SXM4 listed at $1.50 an hour that day, the A100 PCIe at $1.475 per GPU-hour and a 1x H200 NVL at $3.43 (today $1.50/GPU-hr, $1.48/GPU-hr and the console price). The H200 NVL is cheaper per job whenever it finishes at least 2.29 times as fast as an A100 SXM4 (our calculation: 3.43 / 1.50 = 2.29). For BF16 generation, which NVIDIA describes as memory-bound, the paper ratio is 2.35, close to even. For BF16 prefill and training math it is 2.68, and for FP8 serving 5.4, both well past the bar. The A100 wins every hour in which the GPU does little. The A100 SXM4 came in 1 and 8-GPU VMs in us-east-1, us-midwest-1 and us-midwest-2, the A100 PCIe in 2 and 4-GPU VMs in us-midwest-2, and the H200 NVL in 1 and 2-GPU VMs in us-east-1, where a 4-GPU VM appeared 13 minutes after the snapshot. The A100 page and the H200 page list them all, and the pricing page covers billing.
When the H200 NVL is worth the upgrade#
Context is the first case, before a model even outgrows 80 GB. The AWQ build of Llama-3.3-70B, 39.8 GB, serves well on an A100 with about 120,000 tokens of KV cache, and on an H200 NVL the same file leaves about 302,000, 2.5 times as much (our calculation at 320 KiB per token). The FP8 build is where the A100 runs short: its 72.7 GB checkpoint leaves about 5.9 GiB of the A100's 73.6 GiB budget, about 19,000 tokens, and runs weight-only, while one H200 NVL holds it with a full 128k-token sequence of BF16 KV cache, 115.6 GB in all.
Serving to many users is the second. A busy server spends its time in batched matrix math, and with an FP8 checkpoint the H200 NVL does that math at 5.4 times the A100's rate, far past the 2.29 price ratio. It also holds more requests at once: gpt-oss-120b leaves room for 15.2 full 131k-token sequences on the H200 NVL and 2.8 on the A100. vLLM 0.30.0 sets its batch defaults along the same line. On the H200 NVL it allows 1,024 concurrent sequences and 8,192 batched tokens per step for its API server. On any GPU whose name contains "a100" it uses 256 and 2,048, and a comment in its code says a large batched-token limit reduces A100 throughput.
Fine-tuning with room to spare is the third. 16-bit LoRA on a 32B model needs 76 GB in Unsloth's table and 64 to 80 GB in Axolotl's, tight on an A100 and comfortable on an H200 NVL. QLoRA on a 70B model, 41 GB in Unsloth's table, works on both, and the H200 NVL's extra memory goes to longer sequences and bigger batches. The fine-tuning VRAM guide covers the other sizes.
Double precision is the fourth: 30 TFLOPS of FP64 against 9.7, 3.1 times the rate for 2.29 times the price.
When the A100 is still the better buy#
An idle hour is where the A100 holds its ground. Per token of BF16 generation the two cost about the same on paper, so what the A100 saves is the hour the GPU spends waiting: a private assistant a few people use, a notebook left open between experiments, or a QLoRA run on a 70B model, 41 GB in Unsloth's table, that you are still debugging. On 2026-09-27 each of those hours cost $1.50 on an A100 SXM4 against $3.43 on an H200 NVL.
Eight GPUs in one VM is the second. On 2026-09-27 the 8x A100 SXM4 VM put 640 GB of GPU memory in one machine, the most of any configuration in the catalog, for $11.92 an hour, $1.49 per GPU, and QuantaCloud's API flags its GPUs as NVLink-connected. The largest H200 NVL VM listed that day had four GPUs, 564 GB in all, and an NVL bridge joins at most four cards. A job that shards across eight GPUs, such as a full fine-tune with ZeRO-3 or FSDP, stays in the on-demand catalog on the 8x A100. Eight H200s in one server is an HGX H200 or MGX H200 NVL build, which we build to order as reserved capacity: send a capacity brief.
Two A100s in place of one H200 NVL is the third, with a caveat. On 2026-09-27 a 2x A100 PCIe VM listed at $2.95 an hour, 14% less than one H200 NVL, with a 147.2 GiB budget across the pair (our calculation: 2 x 81,920 MiB / 1,024 x 0.92). It holds Llama-3.3-70B at FP8 with a full 128k context, but runs the checkpoint weight-only and splits it across two cards. The single H200 NVL still reads memory faster than both A100 PCIe cards together: 4,800 GB/s against 2 x 1,935 = 3,870.
What fits in 80 GB and in 141 GB#
Nearly 1.8 times the memory changes both which models fit whole and how much context comes with them. The budgets are 73.6 GiB on the A100 and 129.2 GiB on the H200 NVL, at vLLM's default of reserving 92% of GPU memory (our calculation from the 81,920 and 143,771 MiB that nvidia-smi reports).
| Workload | Memory it needs | Source | A100, 80 GB | H200 NVL, 141 GB |
|---|---|---|---|---|
| Llama-3.3-70B, 4-bit AWQ | 39.8 GB checkpoint | Hugging Face | Fits, about 120,000 tokens of KV cache | Fits, about 302,000 tokens |
| gpt-oss-120b | About 65 GB, plus 4.5 GiB per full 131k-token sequence | OpenAI, our calculation | Fits, 2.8 full sequences | Fits, 15.2 full sequences |
| Qwen3-32B, BF16, 32k context | 65.5 GB, plus 8 GiB per sequence | Our calculation | Just inside the budget for one sequence, 4.6 GiB to spare | Fits, 8.5 sequences |
| Llama-3.3-70B, FP8, one full 128k sequence | 72.7 + 42.9 = 115.6 GB | Hugging Face, our calculation | Does not fit on one GPU | Fits |
| Llama-3.3-70B, BF16 | 141.1 GB of weights | Our calculation | 4x A100 PCIe (294.4 GiB budget) or 8x SXM4 | 2x H200 NVL |
| QLoRA, 70B model | 41 GB (Unsloth), 40 to 48 GB (Axolotl) | Unsloth, Axolotl | Fits | Fits |
| 16-bit LoRA, 32B model | 76 GB (Unsloth), 64 to 80 GB (Axolotl) | Unsloth, Axolotl | Tight | Fits |
| 16-bit LoRA, 70B model | 164 GB | Unsloth | 4x A100 PCIe or 8x SXM4 | 2x H200 NVL |
| Wan 2.2 A14B video, 720p | 59.8 GB peak (official script with offloading) | Wan-AI | Fits | Fits |
Every KV figure assumes a BF16 cache and leaves out activations, so treat it as a ceiling. The VRAM guide covers models not listed here, and the KV cache guide works through the per-token formula.
FAQ#
What does the A100 do with an FP8 checkpoint?
It loads it weight-only, because FP8 Tensor Core math needs compute capability 8.9 and the A100 is 8.0. vLLM runs the checkpoint as W8A16: the weights take half the memory of BF16, and the multiplications run at BF16 rates. The H200 NVL runs both per-channel and block-scaled FP8 checkpoints with FP8 math, which is where its 5.4 times paper lead comes from.
Is NVLink working on the multi-GPU VMs?
The API flags both the 8x A100 SXM4 and the H200 NVL offers as NVLink, and nvidia-smi topo -m on the VM is the check. NV followed by a number between two GPUs means NVLink, while PHB, NODE or SYS mean the path runs over PCIe. The InfiniBand vs NVLink guide explains what each link carries.
Should I buy one instead of renting?
That turns on how many hours a year the GPU would be busy. The A100 price guide and the H200 price guide work through purchase prices against hourly rates.
What happens to my data when I stop?
Stopping terminates the VM and deletes its local disk, which on 2026-09-27 was 625 to 1,000 GB on a 1x A100 SXM4 and 750 GB on a 1x H200 NVL, and there are no volumes or snapshots. On a multi-hour fine-tune, copy checkpoints off the VM as you go. Billing stops with the VM, and the unused seconds of the current hour go back to your balance (pricing).
The honest answer is that the upgrade is worth it for busy hours and big models, and not for the rest. If the job fits in about 74 GiB, runs in BF16 and the GPU waits on people, stay on the A100. If it needs more memory, FP8 math, or a GPU that is busy most of the hour, the H200 NVL pays for itself, and for eight GPUs training together the 8x A100 SXM4 is still the on-demand option. RTX PRO 6000 vs H200 covers the 96 GB card in between, L40S vs A100 covers the 48 GB card with FP8, and H200 vs B200 covers the step after the H200.
Launch an H200 NVL Launch an A100 80GB