The difference that matters is memory, not speed per user. The H100 PCIe reads its memory at 2.0 TB/s and the RTX 5090 at 1.79 TB/s. For one person generating tokens, bandwidth sets the ceiling, so on a model that fits in 32 GB the H100's ceiling is only 12 percent higher (our calculation: 2,000 / 1,792 = 1.12). The H100 vs 5090 gap opens everywhere else: 80 GB against 32, 3.6 times the dense FP8 Tensor Core rate with FP32 accumulate (1,513 TFLOPS against 419), and serving software that treats the H100 as a card for many users at once. I would buy the 5090 for my own work under about 29 GB, and rent the H100 by the hour for models that need 30 to 74 GB and for FP8 serving to a team.
H100 PCIe and RTX 5090 specs#
The two cards are closest on bandwidth and furthest apart on memory. The H100 that QuantaCloud rents by the hour is the PCIe card, not the SXM module, so that is the one in the table.
| Spec | H100 PCIe | GeForce RTX 5090 |
|---|---|---|
| Architecture (compute capability) | Hopper (9.0) | Blackwell (12.0) |
| GPU memory | 80 GB HBM2e | 32 GB GDDR7 |
| Memory bandwidth | 2,000 GB/s | 1,792 GB/s |
| FP8 Tensor Core, dense, FP32 accumulate | 1,513 TFLOPS | 419 TFLOPS |
| FP8 Tensor Core, dense, FP16 accumulate | 1,513 TFLOPS | 838 TFLOPS |
| BF16 Tensor Core, dense, FP32 accumulate | 756 TFLOPS | 209.5 TFLOPS |
| FP4 Tensor Core, dense | None | 1,676 TFLOPS |
| ECC | SECDED on HBM2e, caches and register files | Built into the GDDR7 dies, single-bit correction |
| Power | 350 W | 575 W |
| Cooling and size | Passive, dual slot, full height and full length, needs server airflow | 2-slot card, 304 x 137 mm (Founders Edition) |
| GPU to GPU | NVLink bridge to one adjacent card, 600 GB/s | None |
| Multi-Instance GPU | Up to 7 x 10 GB | No |
All Tensor Core figures are dense, without sparsity. The two FP8 rows show the gap that decides serving: NVIDIA lists the H100 at the same 1,513 TFLOPS in both accumulate modes, while the 5090 runs FP8 at half rate when it accumulates in FP32. FP4 goes the other way. The 5090 has FP4 Tensor Cores and the H100 has none, so NVFP4 checkpoints run natively only on the 5090, and vLLM falls back to weight-only 4-bit kernels on the H100. The H100 figures come from NVIDIA's H100 whitepaper and H100 PCIe product brief, and the 5090 figures from its product page and the RTX Blackwell whitepaper. The H100 page compares the PCIe card with the SXM version.
Buying a 5090 against renting an H100#
The break-even is short against the H100: on 2026-09-27 the H100 PCIe had the second highest price per GPU-hour in QuantaCloud's catalog, after the H200 NVL. NVIDIA launched the 5090 at $1,999 in the US on January 30, 2025, and Tom's Hardware reported on 2026-09-05 that the card was "regularly listed for above $5,000" at retail, so check what you would actually pay. On QuantaCloud the H100 PCIe rents from $2.59/GPU-hr (Prices checked 5 Oct 2026, 04:31 UTC). At its price on 2026-09-27, $2.59 an hour, $1,999 buys 772 hours of H100 time and $5,000 buys 1,931 (our calculation: 1,999 / 2.59 and 5,000 / 2.59). At 40 hours a week, that is 19 weeks or 48 weeks.
Per card, the comparison runs the other way. Dated sources put one H100 card at $25,000 to $36,000: Tom's Hardware gave a street price of around $25,000 to $30,000 for the PCIe version in August 2023, CNBC reported in April 2023 that some retailers had offered the H100 at around $36,000, and a distributor listed the PCIe card at $35,315.93, out of stock, on 2026-09-28. That is 12.5 to 18 times the 5090's launch price, and $313 to $450 per GB of memory against $62 for the 5090 (our calculation: 25,000 / 80, 36,000 / 80 and 1,999 / 32). The H100 price guide has the full table.
Each side of the break-even leaves something out. The 5090 still needs a PC with a power supply that can feed a 575 W card, and at that board limit, 772 hours is up to 444 kWh (0.575 kW x 772 hours). The H100's hourly price includes a VM: on 2026-09-27 the 1x H100 PCIe came with 20 vCPUs, 128 GB of RAM and 1,250 GB of disk, in us-midwest-2. Stopping it terminates the VM and deletes that disk, so copy results off first. The unused seconds of the current hour are refunded when you stop, as the pricing page explains.
Launch an H100 PCIeWhere only the H100 will do#
Choose the H100 when the model or the number of users outgrows 32 GB.
Serving an FP8 model to a team is the case the H100 PCIe handles best. Qwen3-32B's FP8 checkpoints are 34.3 GB, more than a 5090 holds at all. On one H100 they leave room for 5.2 full 32k-token conversations at once (our calculation: 73.3 GiB of budget minus 32.0, divided by 8 GiB per conversation). vLLM's own defaults treat the two cards differently too: on a GPU with at least 70 GiB of memory, A100s excepted, vLLM 0.30.0 sets --max-num-seqs to 1,024 concurrent sequences by default, against 256 on smaller cards.
Models that their makers size for 80 GB are the second case. OpenAI sized gpt-oss-120b to "fit into a single 80GB GPU", Google says Gemma 4 31B in BF16 will "fit efficiently on a single 80GB NVIDIA H100 GPU", and Black Forest Labs' reference script for FLUX.2 [dev], a model under a non-commercial license, asks for an "H100-equivalent GPU".
Fine-tuning a 70B model with QLoRA is the third. It needs 40 to 48 GB, and Unsloth's benchmark reached an 89,389-token context for Llama 3.3 70B on 80 GB. A 5090 cannot start that job. The fine-tuning VRAM guide has the sizes for other models.
Where the 5090 matches or beats the H100#
Choose the 5090 when one person uses it every day and the model fits in about 29 GB.
Personal inference is the 5090's strongest case against the H100, because the bandwidth gap is small. gpt-oss-20b runs "within 16GB of memory". Mistral says Devstral Small 2, a 24B model, runs on a "single RTX 4090 or a Mac with 32GB RAM", and a 4090 has 8 GB less than a 5090. For one user on models like these, renting 80 GB buys room you do not use.
NVFP4 models are the second case, and here the H100 cannot follow. Black Forest Labs ships NVFP4 files for FLUX.1 [dev] (9.19 GB) and FLUX.2 [dev] (21.04 GB), and its low-memory setup for FLUX.2 [dev], a 4-bit transformer and text encoder with CPU offload, runs in about 20 GB, a budget BFL aims at RTX 4090 and 5090 cards. ComfyUI computes NVFP4 only on compute capability 10 and up, with a CUDA 13 build of PyTorch.
A fixed budget is the third. A card is a known cost with no meter, and 772 hours of H100 time goes quickly if the GPU is busy every working day. One caution before you buy: Hopper has been supported since CUDA 11.8, while Blackwell needs CUDA 12.8 or newer, so check that every tool you rely on ships sm_120 builds. A100 vs RTX 5090 covers the lower-priced 80 GB rental if you do not need FP8. If you want Blackwell without buying a card, RTX PRO 6000 vs RTX 5090 covers renting one with 96 GB.
What 80 GB holds that 32 GB cannot#
The working budget is 29.3 GiB on the 5090 and 73.3 GiB on the H100, at vLLM's default of 92 percent of GPU memory (our calculation from the 32,607 and 81,559 MiB that nvidia-smi reports).
| Workload | Memory it needs | RTX 5090 | H100 PCIe | Basis |
|---|---|---|---|---|
| gpt-oss-20b | "within 16GB of memory" | Yes | Yes | Published (OpenAI) |
| Gemma 4 31B | 23.3 GB for Google's 4-bit (w4a16) build, 62.5 GB in BF16 | The 4-bit build only, with about 7.6 GiB left for KV cache | Both, BF16 with 15 GiB to spare | Published file sizes and Google's statement, our calculation |
| Qwen3-32B at FP8, 32k-token conversations | 34.3 GB checkpoint, 8 GiB of KV cache per conversation | No | Yes, 5.2 at once | Published file size, our calculation |
| gpt-oss-120b | 65.2 GB of weights | No | Yes, 2.8 full 131,072-token sequences | Published, our calculation |
| Llama-3.3-70B at FP8 | 72.7 GB checkpoint | No | Only at short contexts: about 5.6 GiB left for KV cache, about 18,000 tokens. For long ones use 2x H100 PCIe or one H200 NVL | Published file size, our calculation |
| FLUX.2 [dev] | 35.46 GB fp8 transformer plus 18.03 GB fp8 text encoder | With BFL's 4-bit files and CPU offload, about 20 GB | Yes, both fp8 files on the card | Published (Black Forest Labs, Comfy) |
| LTX-2.5 video | "a minimum 32GB+ VRAM", A100 80 GB or H100 recommended | At the minimum | Recommended | Published (Lightricks) |
| QLoRA, 70B model | 40 to 48 GB (Axolotl), 41 GB (Unsloth) | No | Yes, up to an 89,389-token context (Unsloth) | Published |
The one row where the H100 falls short is a reminder that 80 GB has limits too. A 70B model at FP8 needs two H100s or one H200 NVL with 141 GB for long contexts. The KV cache explainer shows how the conversation counts are worked out.
FAQ#
Is the RTX 5090 faster than the H100?
For FP4, yes, because the H100 has no FP4 Tensor Cores. For FP8 and BF16 with FP32 accumulate, the H100 PCIe has 3.6 times the 5090's dense rate. For one user generating tokens from a model that fits in 32 GB, the two are within 12 percent on memory bandwidth, and that is what sets the speed.
Can I build a server of RTX 5090s instead of renting H100s?
Read NVIDIA's driver license first. Section 2.8 of the NVIDIA Driver License Agreement says GeForce software "is not licensed for datacenter deployment". The 5090 also has no NVLink, so the cards in one server talk over PCIe, and each draws up to 575 W. For eight H100s in one server, we build HGX H100 systems to order as reserved capacity, and the H100 page explains how that works.
Is NVLink enabled on the 2x H100 VM?
QuantaCloud lists its 2x H100 offers as PCIe. The H100 PCIe supports a bridge to one neighbor. Unless the nvidia-smi topo -m entry between the two GPUs starts with NV, plan for PCIe between them, where vLLM recommends pipeline parallelism.
What happens to my data when I stop the H100?
It is deleted. Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy results off first. The unused seconds of the current hour are refunded, and the pricing page covers what is charged and when.
The rule I follow: buy the 5090 for one person's daily work that fits in about 29 GB, and rent the H100 when the model needs 30 to 74 GB, when FP8 serving has to carry many users, or when you need that much GPU for fewer hours a year than the break-even. If the model outgrows 74 GB, I would step up to the H200 NVL before splitting it across two H100s. Launch an H100 PCIe