The short answer: buy the RTX 5090 if your models fit in 32 GB and you will run them most days, and rent the A100 80GB when they do not fit. The A100 vs 5090 choice is mostly about memory. The A100 holds 2.5 times as much: enough for gpt-oss-120b, a 70B model at 4-bit or a QLoRA run on a 70B model, none of which fits on one 5090. On paper the older A100 is not the slow card either: it has more memory bandwidth, 2,039 GB/s on the SXM4 against 1,792, and 1.5 times the 5090's BF16 Tensor Core rate with FP32 accumulate. What the 5090 adds is FP8 and FP4 Tensor Cores, which the A100 lacks, and ownership. At its $1,999 launch price it costs what 1,333 hours of A100 time cost on QuantaCloud on 2026-09-27 (our calculation, below).
A100 80GB and RTX 5090 specs side by side#
The spec sheets split along two lines: memory, where the A100 leads, and number formats, where the 5090 does. QuantaCloud rents the A100 80GB as an SXM4 module and as a PCIe card, so the table gives both where they differ.
| Spec | A100 80GB (SXM4 or PCIe) | GeForce RTX 5090 |
|---|---|---|
| Architecture (compute capability) | Ampere (8.0) | Blackwell (12.0) |
| GPU memory | 80 GB HBM2e | 32 GB GDDR7 |
| Memory bandwidth | 2,039 GB/s SXM4, 1,935 GB/s PCIe | 1,792 GB/s |
| ECC | SECDED on HBM, L2, L1 and register files | Built into the GDDR7 dies, single-bit correction, always on |
| BF16 Tensor Core, dense, FP32 accumulate | 312 TFLOPS | 209.5 TFLOPS |
| FP8 Tensor Core, dense, FP32 accumulate | None | 419 TFLOPS |
| FP4 Tensor Core, dense | None | 1,676 TFLOPS |
| Max power | 400 W SXM4, 300 W PCIe | 575 W |
| Form factor | SXM4 module on an HGX board, or a passive dual-slot PCIe card | 2-slot card, 304 x 137 mm (Founders Edition) |
| GPU to GPU | NVLink at 600 GB/s: SXM4 through the board, PCIe through a 2-card bridge | None, PCIe Gen 5 only |
| Multi-Instance GPU | Up to 7 x 10 GB | No |
| How you get one | Rent it by the hour on QuantaCloud | Buy it: $1,999 at launch |
Every Tensor Core figure above is dense, meaning without sparsity. NVIDIA's headline figures use sparsity and are twice as high, which is where the 5090's 3,352 "AI TOPS" comes from: it is FP4 with sparsity. The figures also use FP32 accumulate, and that choice matters for the 5090 in particular. NVIDIA's whitepaper lists its FP16 math at 419 dense TFLOPS with FP16 accumulate and at half that with FP32 accumulate, and its FP8 at 838 and 419 in the same two modes. The A100 runs FP16 at 312 in both modes. The A100 figures come from NVIDIA's A100 datasheet and Ampere whitepaper, and the 5090 figures from its product page and the RTX Blackwell whitepaper.
Two rows are easy to misread. Both cards have ECC, but not the same ECC: the A100 corrects single-bit errors and detects double-bit errors across its memory, caches and register files, while the 5090's correction lives inside the GDDR7 dies. The NVLink row describes the hardware. Whether a multi-GPU VM exposes it shows only in nvidia-smi topo -m, and the A100 page covers that check for QuantaCloud's 8x SXM4 VM. The 5090 has neither NVLink nor MIG, so a second 5090 talks over PCIe, and one 5090 cannot be split into isolated slices the way an A100 can.
Buying a 5090 against renting an A100 by the hour#
The break-even is 1,333 hours of A100 time at the 5090's launch price. NVIDIA announced the RTX 5090 on January 6, 2025 at $1,999 in the US, on sale from January 30. Street prices have run far above that: Tom's Hardware reported on 2026-09-05 that the card was "regularly listed for above $5,000". On QuantaCloud the A100 80GB rents from $1.48/GPU-hr (Prices checked 6 Oct 2026, 02:10 UTC). The arithmetic uses the price on 2026-09-27, $1.50 an hour for a 1x A100 SXM4.
| Price paid for the 5090 | A100 hours it buys at $1.50 | At 10 hours a week | At 40 hours a week |
|---|---|---|---|
| $1,999, NVIDIA's launch price | 1,333 | 133 weeks | 33 weeks |
| $5,000, the street level Tom's Hardware reported | 3,333 | 333 weeks | 83 weeks |
Our calculation: the card's price divided by $1.50, then by the hours per week. Both sides leave something out. The 5090 needs a PC around it: a CPU, motherboard, RAM, storage and a power supply that can feed the card. At its 575 W board limit, 1,333 hours is up to 766 kWh, billed at your local electricity rate (0.575 kW x 1,333 hours). A card you own keeps some resale value, which the table ignores. The A100's hourly price covers a whole VM: on 2026-09-27 the 1x A100 SXM4 came with 14 to 16 vCPUs, 100 to 120 GB of RAM and 625 to 1,000 GB of disk. What renting does not give you is a disk that survives. Stopping an instance terminates it and deletes its disk, there are no volumes or snapshots, and results have to be copied off first. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. The pricing page has the full rules.
Buying an A100 is not the realistic alternative. New A100 80GB PCIe cards were listed at $13,999.95 and $16,535.00 on 2026-09-28, both out of stock, seven to eight times the 5090's launch price. The A100 price guide has the dated sources.
Launch an A100 80GBWhen the A100's 80 GB is worth renting#
Choose the A100 when the job needs more than 32 GB on one GPU, or several 80 GB GPUs in one machine.
Large-model inference is where the gap is widest. OpenAI sized gpt-oss-120b to "fit into a single 80GB GPU", its MXFP4 weights take 65.2 GB, and vLLM runs MXFP4 on the A100 through its Marlin kernels. No single 5090 can load it. The gpt-oss GPU requirements guide covers the model in detail.
Fine-tuning a 70B model, or a 32B model in 16-bit, is the second. QLoRA on a 70B model needs 41 GB by Unsloth's count and 40 to 48 GB by Axolotl's, which one A100 holds with room left and a 5090 cannot hold at all. 16-bit LoRA on a 32B model needs 64 to 80 GB, tight on one A100. For a full fine-tune, the 8x A100 SXM4 VM that was in the catalog on 2026-09-27 puts 640 GB of GPU memory in one machine. The LoRA, QLoRA and full fine-tuning guide has the figures for each model size.
Video at 720p is the third. Wan 2.2 A14B peaked at 59.8 GB on one GPU in Wan's own 720p test, even with offloading. Lightricks lists 32 GB as the minimum for LTX-2.5 and recommends an A100 80GB or an H100, and LTX-2.5's license is free for commercial use below $10M in annual revenue.
Occasional heavy work is the last. If you need 80 GB for about 20 hours a month, 1,333 hours of A100 time lasts five and a half years (our calculation: 1,333 / 20 = 67 months), and nothing sits idle between jobs.
When a 5090 on your desk is the better buy#
Choose the 5090 when the work fits in 32 GB and you run it most days.
Daily inference on models that fit is the main case. vLLM claims 92 percent of GPU memory by default, which leaves a 29.3 GiB working budget on a 5090 (our calculation: the card reports 32,607 MiB, and 32,607 / 1,024 x 0.92 = 29.3). That covers gpt-oss-20b, which OpenAI says runs "within 16GB of memory", 27B models such as Qwen3.8-27B at 4-bit (at least 13.9 GB of weights by parameter count), and Meta's Muse Glimmer 30B, whose 4-bit model Meta sizes for "a 24 GB or 32 GB envelope". Past the break-even, each extra hour on your own card costs only electricity.
FP8 and FP4 image generation is the second. Black Forest Labs publishes FP8 (12.33 GB) and NVFP4 (9.19 GB) versions of FLUX.1 [dev], a model under a non-commercial license. ComfyUI computes in FP8 only on compute capability 8.9 and up, and in NVFP4 only on 10 and up with a CUDA 13 build of PyTorch. The 5090 computes in both formats on its Tensor Cores, and the A100 in neither.
Local development is the third. A card in your own machine keeps its disk, your data stays with you, and no meter runs while you read a stack trace. Budget time for the software once: Blackwell needs CUDA 12.8 or newer and a PyTorch build with sm_120 kernels, and older environments stop with no kernel image is available for execution on the device. The driver and CUDA version guide shows how to check what you have, and if your models fit in 48 GB, RTX A6000 vs RTX 5090 compares the cheaper rental.
What fits in 32 GB, and what needs 80#
The working budgets are 29.3 GiB on the 5090 and 73.6 GiB on the A100, at vLLM's default of 92 percent (our calculation from the 32,607 and 81,920 MiB that nvidia-smi reports). Weights, KV cache and activations share that budget.
| Workload | Memory it needs | RTX 5090 | A100 80GB | Basis |
|---|---|---|---|---|
| gpt-oss-20b | "within 16GB of memory" | Yes | Yes | Published (OpenAI) |
| Qwen3-8B in BF16, 32k-token conversations | 16.4 GB (15.3 GiB) of weights, 4.5 GiB of KV cache per conversation | 3.1 at once: (29.3 - 15.3) / 4.5 | 13.0 at once: (73.6 - 15.3) / 4.5 | Our calculation |
| Qwen3.8-27B in BF16 | 55.6 GB checkpoint, 64 KiB of KV cache per token | No. At 4-bit (at least 13.9 GB by parameter count), yes | Yes, with 21.8 GiB left, room for one full 262,144-token context (16 GiB) before activations | Published file size, our calculation |
| Llama-3.3-70B at 4-bit (AWQ build) | 39.8 GB checkpoint | No | Yes, about 120,000 tokens of BF16 KV cache | Published file size, our calculation |
| gpt-oss-120b | 65.2 GB of weights, 4.5 GiB of KV per full 131,072-token sequence | No | Yes, 2.8 full-length sequences | Published, our calculation |
| 16-bit LoRA, 8B model | 22 GB (Unsloth), 16 to 24 GB (Axolotl) | Yes | Yes | Published |
| QLoRA, 32B model | 26 GB (Unsloth), 24 to 32 GB (Axolotl) | Tight | Yes | Published |
| QLoRA, 70B model | 41 GB (Unsloth), 40 to 48 GB (Axolotl) | No | Yes | Published |
| Full fine-tune, 7B to 8B, BF16 with AdamW | 60 to 80 GB (Axolotl) | No | Tight on one, comfortable across the 8x SXM4 VM | Published |
| Wan 2.2 A14B video at 720p | 59.8 GB peak on one GPU, with offloading | No | Yes | Published (Wan) |
ComfyUI can offload model weights to system RAM, so the image and video rows bend on both cards at the cost of speed. For models that are not in the table, the VRAM guide and the KV cache explainer show the arithmetic.
FAQ#
Is the RTX 5090 faster than the A100?
For FP8 and FP4 work, on paper yes: the 5090 has Tensor Cores for both, and the A100 has neither. For BF16 math with FP32 accumulate the A100 is ahead, 312 dense TFLOPS to 209.5, and it has more memory bandwidth, which is what sets the pace of token generation for a single user. None of that matters once a model needs more than 29 GB, because the 5090 cannot hold it on one card.
Can two RTX 5090s replace one A100 80GB?
Not fully. Two 5090s hold 64 GB, still less than 80, and they talk over PCIe because the 5090 has no NVLink. vLLM's docs recommend pipeline parallelism rather than tensor parallelism on GPUs without NVLink. Two cards also draw up to 1,150 W at their board limits (our calculation: 2 x 575 W), before the rest of the machine.
Does QuantaCloud rent the RTX 5090?
No. QuantaCloud rents data-center and workstation GPUs by the hour, from the RTX A6000 to the H200 NVL, and the GPU catalog lists all of them with live prices. If you need FP8 and 80 GB together, H100 vs RTX 5090 is the comparison to read next. RTX PRO 6000 vs RTX 5090 covers renting a 96 GB Blackwell card instead.
What happens to my files when I stop an A100?
They are deleted with the instance. Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy checkpoints off while a fine-tune runs. The unused seconds of the current hour are refunded, and the pricing page explains what is charged and when.
My rule for this pair: if the model and its context fit in about 29 GB and you will use the card most days, buy the 5090. If the job needs more than that, or you need that much memory for fewer hours than the break-even above, rent the A100 and stop it when the job ends. Start on one A100 to confirm the job fits, then move the same code to the 8x VM if it needs more. Launch an A100 80GB