GPU comparison

RTX A6000 vs RTX 5090

A6000 vs 5090 for AI: the RTX 5090 is faster, the RTX A6000 holds 16 GB more. What fits in 48 GB but not 32 GB, and when renting beats buying.

Faiz Ahmed11 min read
NVIDIA · Ampere

RTX A6000

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
768 GB/s
FP8 compute
No native FP8
From / GPU-hr$0.48
Launch RTX A6000
NVIDIA · Blackwell

RTX 5090

Hardware for your own workstation

32 GBGDDR7 memory
Bandwidth
1,792 GB/s
FP8 compute
Native FP8
How you get itRetail hardware
Read the comparison

Prices checked

The honest answer is that the RTX 5090 is the faster card and the RTX A6000 is the bigger one. The 5090 reads memory 2.3 times as fast, 1,792 GB/s against 768, and has FP8 and FP4 Tensor Cores that the Ampere-based A6000 lacks. The A6000 has 48 GB against 32, and those 16 GB decide whether a 70B model at 4-bit, a QLoRA run on a 70B model or Qwen-Image in BF16 fits on one card at all. If your work fits in about 29 GB and you run it most days, buy the 5090. If it needs 30 to 44 GB, or you only need a GPU some evenings, the A6000 vs 5090 arithmetic favors renting the A6000 by the hour.

RTX A6000 and RTX 5090 specs#

Both cards plug into an ordinary PCIe slot, and they sit two architectures apart: Ampere, then Ada, then Blackwell. The A6000 is NVIDIA's Ampere workstation card, the 5090 its Blackwell GeForce flagship.

SpecRTX A6000GeForce RTX 5090
Architecture (compute capability)Ampere GA102 (8.6)Blackwell GB202 (12.0)
GPU memory48 GB GDDR6 with ECC32 GB GDDR7, ECC built into the memory dies
Memory bandwidth768 GB/s1,792 GB/s
CUDA cores10,75221,760
BF16 Tensor Core, dense, FP32 accumulate154.8 TFLOPS209.5 TFLOPS
FP16 Tensor Core, dense, FP16 accumulate154.8 TFLOPS419 TFLOPS
FP8 Tensor Core, denseNone419 TFLOPS with FP32 accumulate, 838 with FP16
FP4 Tensor Core, denseNone1,676 TFLOPS
Power300 W575 W
Size and coolingDual slot, 4.4 x 10.5 in, active2-slot, 304 x 137 mm (Founders Edition)
NVLink2-way bridge, 112.5 GB/sNone
How you get oneRent it by the hour on QuantaCloudBuy it: $1,999 at launch

All Tensor Core figures are dense, without sparsity. The accumulate rows show a real difference between a workstation card and a GeForce card: NVIDIA lists the A6000 at the same 154.8 TFLOPS whether FP16 math accumulates in FP16 or FP32, and the 5090 at 419 with FP16 accumulate but 209.5 with FP32. Even in its slower mode the 5090 does 1.35 times the A6000's BF16 math (our calculation: 209.5 / 154.8). The A6000 figures are from NVIDIA's A6000 datasheet and the dense table in its Ada professional whitepaper, and the 5090 figures from its product page and the RTX Blackwell whitepaper.

ECC is on both lists, in different forms. The A6000's GDDR6 has ECC, and NVIDIA says the 5090's ECC is built into the GDDR7 dies, always on, with single-bit correction. The A6000 also supports a 2-way NVLink bridge, which the 5090 does not, although whether a QuantaCloud multi-GPU VM has a bridge shows only in nvidia-smi topo -m. The RTX A6000 page has the card's full specs, and RTX A6000 vs RTX 4090 compares it with the RTX 4090.

Buy a 5090 or rent the A6000 by the hour#

The A6000 is the hardest rental for a 5090 to beat on cost. On QuantaCloud it rents from $0.48/GPU-hr (Prices checked 5 Oct 2026, 22:53 UTC), and on 2026-09-27 the lowest 1x price was $0.48 an hour. At that price the $1,999 NVIDIA asked when the 5090 went on sale on January 30, 2025 equals 4,165 hours of A6000 time, and $5,000, the street level Tom's Hardware reported on 2026-09-05 ("regularly listed for above $5,000"), equals 10,417 hours (our calculation: 1,999 / 0.48 and 5,000 / 0.48). This is how long those hours last:

Hours a week on the GPUWeeks to break even at $1,999Weeks to break even at $5,000
10416, about 8 years1,042, about 20 years
20208, about 4 years521, about 10 years
40104, about 2 years260, about 5 years

Our calculation: break-even hours divided by hours per week, and weeks divided by 52. The purchase side leaves out the PC around the card, a power supply that can feed it, and electricity: at the 575 W board limit, 4,165 hours is up to 2,395 kWh (0.575 kW x 4,165 hours). The rental side includes a VM: on 2026-09-27 a 1x A6000 came with 6 to 12 vCPUs, 24 to 64 GB of RAM and 256 GB of disk, and a 2x with 96 GB of GPU memory started at $0.96 an hour. What the rental does not keep is data. Stopping an instance terminates it and deletes its disk, so copy results off first. The pricing page covers the billing rules.

Launch an RTX A6000

What the A6000's extra 16 GB buys#

Choose the A6000 when the job needs 30 to 44 GB on one card.

A 70B model at 4-bit is the defining case. vLLM claims 92 percent of GPU memory by default, a 41.4 GiB budget on the A6000 with ECC on, when it reports 46,068 MiB. Llama-3.3-70B comes to 35.3 GB at 4-bit on paper, but the AWQ build used below is a 39.8 GB download, because 4-bit files carry scales and keep the embedding and output layers in 16-bit. That leaves about 4.4 GiB for KV cache, roughly 14,000 tokens at BF16 (our calculation: 4.4 GiB / 320 KiB per token), enough for a long chat but not a long document. With ECC off the card reports 49,140 MiB, which adds 2.8 GiB. The same weights do not load on a 5090 at all.

Fine-tuning that needs the extra 16 GB is the second. QLoRA on a 70B model needs 40 to 48 GB by Axolotl's count and 41 GB by Unsloth's, and Unsloth's benchmark reached a 12,106-token context for Llama 3.3 70B on 48 GB. Axolotl's own Llama-3 8B example runs a full fine-tune on a single 48 GB GPU, with an 8-bit paged AdamW optimizer and gradient checkpointing. Neither job starts on 32 GB. The LoRA, QLoRA and full fine-tuning guide covers the other sizes.

Image and video models at full precision are the third. Qwen-Image's BF16 model file alone is 40.9 GB, and Wan 2.2 A14B peaked at 41.3 GB on one GPU in Wan's 480p test, even with offloading. On a 5090 you drop to the fp8 files or lean on ComfyUI's offloading, which costs speed. The A6000 cannot compute in FP8, so fp8 files save memory on it but not time. The ComfyUI page covers the template.

What the 5090's bandwidth and FP8 buy#

Choose the 5090 when speed matters and the model fits in about 29 GB.

Token generation is where the bandwidth gap shows. NVIDIA describes LLM token generation as memory-bound, so a single stream is capped by how fast the card reads its weights: about 47 tokens per second for Qwen3-8B in BF16 (16.4 GB) on the A6000, and about 109 on the 5090 (our calculation: 768 / 16.4 and 1,792 / 16.4). Those are upper bounds, not measurements, and real runs land below them.

FP8 and FP4 are the second case. FP8 Tensor Core math needs compute capability 8.9, and the A6000 is 8.6, so vLLM runs FP8 checkpoints on it weight-only (W8A16) and NVFP4 as weight-only 4-bit, and ComfyUI will not compute in FP8 on it. On the 5090 both formats run on Tensor Cores built for them: 419 dense FP8 TFLOPS with FP32 accumulate and 1,676 for FP4.

Daily use at a desk is the third. Once the card is paid for, an hour of use costs only electricity, and your files stay on your own disk between sessions.

What fits in 48 GB but not in 32 GB#

The working budgets are 41.4 GiB on the A6000 and 29.3 GiB on the 5090, at vLLM's default of 92 percent (our calculation from the 46,068 and 32,607 MiB that nvidia-smi reports).

WorkloadMemory it needsRTX 5090RTX A6000Basis
Qwen3-8B in BF16, 32k-token conversations16.4 GB of weights, 4.5 GiB of KV per conversation3.1 at once5.8 at onceOur calculation
16-bit LoRA, 14B model33 GB (Unsloth), 28 to 40 GB (Axolotl)NoYesPublished
QLoRA, 32B model26 GB (Unsloth), 24 to 32 GB (Axolotl)TightYesPublished
QLoRA, 70B model41 GB (Unsloth), 40 to 48 GB (Axolotl)NoTight, 12,106-token context (Unsloth)Published
Full fine-tune, Llama-3 8BOne 48 GB GPU with 8-bit paged AdamW and gradient checkpointing (Axolotl example)NoYesPublished
Llama-3.3-70B at 4-bit (AWQ build)39.8 GB checkpointNoYes, about 14,000 tokens of BF16 KVPublished file size, our calculation
Qwen3.6-35B-A3B, Qwen's FP8 checkpoint37.5 GBNo. At 4-bit, at least 18 GB by parameter count, yesYes, weight-only FP8, 6.5 GiB leftPublished file size, our calculation
Qwen-Image, BF16 model file40.9 GB, before the 9.4 GB fp8 text encoderNo. The 20.4 GB fp8 file, yesYes, with the text encoder offloaded to system RAMPublished file sizes
Wan 2.2 A14B video at 480p41.3 GB peak on one GPU, with offloadingNot at the published figureYesPublished (Wan)
gpt-oss-120b65.2 GB of weightsNoNo. A 2x A6000 VM holds it with about 22 GiB to sparePublished, our calculation

The last row is the limit of both cards. For models that are not in the table, the VRAM guide shows the arithmetic, and fixing CUDA out of memory covers what to try before you move to a bigger card.

Older CUDA stacks run on the A6000, and Blackwell needs new ones#

The software gap runs the other way from the speed gap. The A6000 is compute capability 8.6, and every current CUDA 12 and 13 release supports it, as does every current PyTorch build, the cu126 wheels included. The 5090 is compute capability 12.0 (sm_120), which needs CUDA 12.8 or newer, an R570 or newer driver and, on Linux, NVIDIA's open GPU kernel modules. PyTorch 2.7.0, released on April 23, 2025, was the first version with Blackwell kernels, and only in its CUDA 12.8 wheels. A build without them stops with with CUDA capability sm_120 is not compatible with the current PyTorch installation or no kernel image is available for execution on the device.

The fix is a newer wheel: cu128 for PyTorch 2.7 to 2.11, cu130 from 2.9 or cu132 from 2.12. Plain pip install torch on Linux has installed CUDA 13.0 wheels since PyTorch 2.11, and those need an R580 or newer driver. This one line shows whether a build includes Blackwell kernels:

Terminal
python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.get_arch_list())"

On a build that supports the 5090, the list includes sm_120. The driver and CUDA version guide covers the other checks. To try your stack on Blackwell before you buy a card, rent the RTX PRO 6000 Blackwell on QuantaCloud: it has the same compute capability, 12.0. RTX PRO 6000 vs RTX 5090 compares the two cards.

FAQ#

Is 32 GB enough for local AI work?

For a lot of it. SDXL runs on 8 GB cards, FLUX.1 [dev] in FP8 is a 12.33 GB file, gpt-oss-20b runs within 16 GB, and 27B to 32B models run at 4-bit. What 32 GB cannot hold on one card is a 70B model, QLoRA on a 70B model, or an image or video model whose BF16 file passes 30 GB.

Does the RTX A6000 support FP8?

No. FP8 Tensor Core math needs compute capability 8.9, and the A6000 is 8.6. FP8 files still load and halve the weight memory, but the math runs in 16-bit. If you want FP8 speed at 48 GB, RTX 6000 Ada vs RTX 5090 covers the Ada card with the same memory.

Can I rent two or four A6000s in one VM?

Yes. On 2026-09-27 the catalog offered 1, 2 or 4 per VM, with 48, 96 or 192 GB of GPU memory. QuantaCloud does not claim NVLink between them, and without it vLLM recommends pipeline parallelism over tensor parallelism. The RTX A6000 page has the live list.

What happens to my files when I stop?

They are deleted. Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy results off first. The unused seconds of the current hour are refunded, as the pricing page explains.


My rule for these two: if the job fits in about 29 GB and runs most days, buy the 5090. If it needs 30 to 44 GB, or it runs a few evenings a week, rent the A6000 and spend the difference on more hours. For image and video work, launch it with the ComfyUI template. Launch an RTX A6000 with ComfyUI

Keep building

Choose your next step.