GPU comparison

RTX A6000 vs RTX 4090

A6000 vs 4090 for AI: twice the memory against 31% more bandwidth and FP8. What 48 GB fits that 24 GB cannot, and renting vs buying a 4090 now.

Faiz Ahmed11 min read
NVIDIA · Ampere

RTX A6000

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
768 GB/s
FP8 compute
No native FP8
From / GPU-hr$0.48
Launch RTX A6000
NVIDIA · Ada Lovelace

RTX 4090

Hardware for your own workstation

24 GBGDDR6X memory
Bandwidth
1,008 GB/s
FP8 compute
Native FP8
How you get itRetail hardware
Read the comparison

Prices checked

The RTX 4090 is the faster card and the RTX A6000 the roomier one, and for most AI work the room decides. The 4090 reads memory 1.31 times as fast, 1,008 GB/s against 768, and has FP8 Tensor Cores that the Ampere-based A6000 lacks, but in BF16, which NVIDIA rates only with FP32 accumulate, it is just 7 percent ahead. The A6000 has 48 GB against 24, and that decides whether a 14B model in BF16, a 70B model at 4-bit or QLoRA on a 32B model fits at all. Tom's Hardware's price tracker also says the 4090 is out of production, and on 2026-09-14 it listed $3,220 as the lowest US price for an available card, about twice the $1,599 launch price. I would buy a 4090 only for daily work that fits in about 22 GB, and rent the A6000 by the hour for anything between 23 and 44 GB.

RTX A6000 and RTX 4090 specs#

The 4090 is one generation newer than the A6000, and the table shows what that generation changed.

SpecRTX A6000GeForce RTX 4090
Architecture (compute capability)Ampere GA102 (8.6)Ada Lovelace AD102 (8.9)
GPU memory48 GB GDDR6 with ECC24 GB GDDR6X, no ECC listed
Memory bandwidth768 GB/s1,008 GB/s
L2 cache6 MB72 MB
CUDA cores10,75216,384
Tensor Cores336, third generation512, fourth generation
BF16 Tensor Core, dense, FP32 accumulate154.8 TFLOPS165.2 TFLOPS
FP16 Tensor Core, dense, FP16 accumulate154.8 TFLOPS330.3 TFLOPS
FP8 Tensor Core, denseNone330.3 TFLOPS with FP32 accumulate, 660.6 with FP16
FP4 Tensor CoreNoneNone
Power300 W450 W
Size and coolingDual slot, 4.4 x 10.5 in, fan3 slots, 304 x 137 mm (Founders Edition)
NVLinkBridge for two cards, 112.5 GB/sNone
How you get oneRent it by the hour on QuantaCloudBuy it: $1,599 at launch in 2022

All Tensor Core figures are dense, without sparsity, and the accumulate rows hide the surprise in this pair. NVIDIA lists the A6000 at 154.8 TFLOPS whether FP16 math accumulates in FP16 or FP32, while the 4090, a GeForce card, runs at half rate with FP32 accumulate. BF16 is listed only with FP32 accumulate, and there the 4090 leads by just 7 percent, 165.2 TFLOPS against 154.8 (our calculation: 165.2 / 154.8 = 1.07). With FP16 accumulate it leads 2.1 to 1 (330.3 / 154.8), and in FP8 it has a format the A6000 does not. The 4090 figures come from NVIDIA's Ada architecture whitepaper and product page, and the A6000 figures from its datasheet and the Ada professional whitepaper.

Two more rows matter for a desk. The 4090 Founders Edition takes three slots, and NVIDIA asks for 850 W of system power, based on a PC with a Ryzen 9 5900X. The A6000's GDDR6 carries ECC, and NVIDIA's spec page for the 4090 lists none. The RTX A6000 page has the card's full specs and live configurations.

A 4090 at twice its launch price, against A6000 hours#

The 4090 now costs about twice its launch price, which moves the break-even out by years. NVIDIA put the RTX 4090 on sale on October 12, 2022, starting at $1,599. Tom's Hardware's GPU price tracker, last updated on 2026-09-14, lists $3,220 as the lowest US price for an available 4090, says the 40 series is no longer produced, and warns that cards for sale have a high chance of being second-hand or ex-mining hardware. On QuantaCloud the A6000 rents from $0.48/GPU-hr, and on 2026-09-27 the lowest 1x price was $0.48 an hour. At that price $1,599 buys 3,331 hours of A6000 time and $3,220 buys 6,708 (our calculation: 1,599 / 0.48 and 3,220 / 0.48).

Prices checked 5 Oct 2026, 21:30 UTC

Hours a week on the GPUWeeks to break even at $1,599Weeks to break even at $3,220
10333, about 6.4 years671, about 12.9 years
20167, about 3.2 years335, about 6.4 years
4083, about 1.6 years168, about 3.2 years

Our calculation: break-even hours divided by hours per week, then weeks divided by 52. Neither side of the table is the whole cost. A 4090 needs a PC with the room and the power supply for a 450 W card, and at that board limit 3,331 hours can draw up to 1,499 kWh (0.45 kW x 3,331 hours). The A6000's hourly price covers a whole VM, which on 2026-09-27 had 6 to 12 vCPUs, 24 to 64 GB of RAM and 256 GB of disk, but nothing on it survives a stop: stopping terminates the instance and deletes its disk, so copy results off first. The unused seconds of the current hour are refunded, as the pricing page explains.

Launch an RTX A6000

Where 24 GB runs out#

Choose the A6000 when the job needs 23 to 44 GB on one card.

Models between 14B and 32B are where 24 GB runs out first. vLLM claims 92 percent of GPU memory by default, a 22.1 GiB budget on a 4090 and 41.4 GiB on the A6000 with ECC on (our calculation from the 24,564 and 46,068 MiB that nvidia-smi reports). Qwen3-14B in BF16 is a 29.5 GB checkpoint, so it does not load on a 4090, and on the A6000 it leaves room for 2.8 full 32k-token conversations at 5 GiB each. Qwen's AWQ 4-bit build of Qwen3-32B, 19.3 GB, loads on both, but on the 4090 it leaves about 4.1 GiB of KV cache, roughly 17,000 tokens, where the A6000 has room for 2.9 full 32k conversations.

A 70B model at 4-bit is the next step. The AWQ build of Llama-3.3-70B is a 39.8 GB download, too big for a 4090, and on the A6000 it leaves about 4.4 GiB of KV cache, roughly 14,000 tokens at 320 KiB per token. With ECC off the A6000 reports 49,140 MiB instead of 46,068, which adds 2.8 GiB.

Fine-tuning is the third case. Unsloth's own benchmark shows QLoRA on Llama 3.1 8B reaching a 78,475-token context on 24 GB and 191,728 on 48 GB. QLoRA on a 32B model needs 26 GB by Unsloth's count, which it calls the absolute minimum and which is more than a 4090 has, and 24 to 32 GB by Axolotl's. QLoRA on a 70B model needs 40 to 48 GB, which the A6000 holds with a 12,106-token context in Unsloth's benchmark. The LoRA, QLoRA and full fine-tuning guide covers the other sizes.

Image and video models at full size are the fourth. Comfy's docs ran Qwen-Image's fp8 files, a 20.4 GB model and a 9.4 GB text encoder, on a 24 GB RTX 4090D at 86 percent of its memory, so 24 GB copes with the fp8 route. The BF16 model file alone is 40.9 GB, and Wan 2.2 A14B peaked at 41.3 GB on one GPU in Wan's own 480p test, with offloading. Both need the A6000, or ComfyUI's offloading on a 4090, which costs speed. The ComfyUI page covers the template.

Where the 4090's speed pays#

Buy the 4090 when the model fits in about 22 GB and speed is what you are paying for.

Single-user chat is where its bandwidth shows. Every generated token rereads the weights, so Qwen3-8B's 16.4 GB of BF16 weights cap one stream at about 61 tokens per second on the 4090 and about 47 on the A6000 (our calculation: 1,008 / 16.4 and 768 / 16.4). Those are ceilings, not measurements.

FP8 is the second case. The 4090 is compute capability 8.9, so vLLM runs per-channel FP8 checkpoints on it with FP8 math, and ComfyUI computes fp8 files, such as Black Forest Labs' 12.33 GB FP8 file of FLUX.1 [dev], in FP8. The A6000 loads the same files and runs the math in 16-bit: FP8 saves it memory, not time. Neither card has FP4, which is Blackwell-only.

Models sized for a 24 GB card are the third. OpenAI says gpt-oss-20b runs "within 16GB of memory", Meta sizes Muse Glimmer 30B at 4-bit for "a 24 GB or 32 GB envelope", and Wan names the RTX 4090 as its example card for running TI2V-5B video at 720p. For work like that, a 4090 you own runs every day for the cost of electricity once the break-even passes, and your files stay on your own disk.

What 48 GB fits that 24 GB cannot#

Every row is checked against the same budgets, 22.1 GiB on the 4090 and 41.4 GiB on the A6000, and the last column says where each figure comes from.

WorkloadMemory it needsRTX 4090RTX A6000Basis
gpt-oss-20b13.8 GB of weights, "within 16GB of memory"YesYesPublished (OpenAI, Hugging Face)
Qwen3-8B in BF16, 32k-token conversations16.4 GB of weights, 4.5 GiB of KV per conversation1.5 at once5.8 at onceOur calculation
Qwen3-14B in BF1629.5 GB checkpoint, 5 GiB of KV per 32k conversationNoYes, 2.8 conversations at oncePublished file size, our calculation
Qwen3-32B, AWQ 4-bit19.3 GB checkpoint, 8 GiB of KV per 32k conversationTight: about 17,000 tokens of KVYes, 2.9 conversations at oncePublished file size, our calculation
Llama-3.3-70B, AWQ 4-bit39.8 GB checkpointNoYes, about 14,000 tokens of KVPublished file size, our calculation
QLoRA on Llama 3.1 8B, longest contextUnsloth's benchmark78,475 tokens191,728 tokensPublished (Unsloth)
QLoRA, 32B model26 GB (Unsloth), 24 to 32 GB (Axolotl)Below Unsloth's minimumYesPublished
QLoRA, 70B model41 GB (Unsloth), 40 to 48 GB (Axolotl)NoTight, 12,106-token context (Unsloth)Published
Qwen-Image, fp8 model and text encoder20.4 GB + 9.4 GB of filesYes, at 86 percent of a 24 GB 4090D in Comfy's testYesPublished (Comfy)
Wan 2.2 A14B video at 480p41.3 GB peak on one GPU, with offloadingNot at the published figureYesPublished (Wan)
gpt-oss-120b65.2 GB of weightsNoNo. A 2x A6000 VM holds it with about 22 GiB to sparePublished, our calculation

gpt-oss-120b is where both single cards stop, and two A6000s in one VM are the rented way past it. The image and video rows are softer limits than the LLM rows, because ComfyUI can offload weights to system RAM at a cost in speed. The VRAM guide and the KV cache explainer show the arithmetic for models not listed here.

FAQ#

Does the RTX 4090 have ECC memory?

NVIDIA does not list it. The 4090's spec page gives 24 GB of GDDR6X with no ECC line, while the RTX A6000's datasheet lists ECC on its 48 GB of GDDR6. I count that as one more point for the workstation card on long training runs.

Can two RTX 4090s replace one RTX A6000?

For some jobs. Two 4090s hold 48 GB between them, but they talk over PCIe, since the 4090 has no NVLink, and for a model split across them vLLM's docs recommend pipeline parallelism over tensor parallelism. Axolotl documents FSDP with QLoRA training a 70B model on two 24 GB GPUs, and Hugging Face's PEFT docs measured that kind of run at 19.6 GB per GPU with CPU offload and about 107 GB of host RAM. Two cards also draw up to 900 W at their board limits (our calculation: 2 x 450 W).

Can I rent an RTX 4090 on QuantaCloud?

No. The 4090 is a GeForce card, and QuantaCloud rents data-center and workstation GPUs, each listed with its live price in the GPU catalog. The closest rental is the RTX 6000 Ada: the 4090's AD102 chip with 48 GB of ECC memory and FP8, compared in RTX 6000 Ada vs RTX 4090.


The rule I use for this pair is the 22 GB line. Under it, with daily use, a 4090 pays for itself, as long as you buy it from a reputable seller now that the card is out of production. Over it, or for a few evenings a week, rent the A6000: at the 2026-09-27 price, 3,331 hours of it cost what a 4090 did at launch. If you want FP8 math at 48 GB, rent the RTX 6000 Ada instead, and RTX A6000 vs RTX 5090 runs the same arithmetic against the newer 32 GB card. Launch an RTX A6000

Keep building

Choose your next step.