GPU comparison

RTX A6000 vs A100

RTX A6000 vs A100 80GB: specs, live hourly prices, what fits in 48 GB against 80 GB, and when the cheaper A6000 does the same job for less.

Faiz Ahmed10 min read
NVIDIA · Ampere

RTX A6000

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
768 GB/s
FP8 compute
No native FP8
From / GPU-hr$0.48
Launch RTX A6000
NVIDIA · Ampere

A100

Cloud GPU with full SSH access

80 GBHBM2e memory
Bandwidth
Up to 2,039 GB/s
FP8 compute
No native FP8
From / GPU-hr$1.50
Launch A100

Prices checked

The short answer is: the RTX A6000 whenever the job fits in 48 GB, the A100 when it does not or when someone is waiting on the result. The A100 is the faster card, with up to 2.7 times the memory bandwidth and twice the BF16 Tensor Core rating. On 2026-09-27 it also cost 3.1 times as much per hour: $1.50 for a 1x A100 SXM4 against $0.48 for a 1x A6000. On paper that makes the A6000 the cheaper way to finish any job it can hold. The A100 earns its price when the job needs 80 GB: a 70B model with a long context, a 32B model in BF16, gpt-oss-120b on one GPU, or 16-bit LoRA on a 30B model.

RTX A6000 and A100 specs side by side#

The spec sheets split cleanly: the A100 has more memory, more bandwidth and more Tensor Core math, and the A6000 has more plain FP32. I compare the A6000 with both 80 GB A100s that QuantaCloud sells, the PCIe card and the SXM4 module.

SpecRTX A6000A100 80GB PCIeA100 80GB SXM4
Architecture (compute capability)Ampere (8.6)Ampere (8.0)Ampere (8.0)
GPU memory48 GB GDDR6 with ECC80 GB HBM2e80 GB HBM2e
Memory bandwidth768 GB/s1,935 GB/s2,039 GB/s
FP3238.7 TFLOPS19.5 TFLOPS19.5 TFLOPS
TF32 Tensor Core (dense)77.4 TFLOPS156 TFLOPS156 TFLOPS
BF16 and FP16 Tensor Core (dense)154.8 TFLOPS312 TFLOPS312 TFLOPS
FP8 Tensor CoreNot supportedNot supportedNot supported
FP64About 0.6 TFLOPS (1/64 of FP32, our calculation)9.7 TFLOPS, 19.5 on Tensor Cores9.7 TFLOPS, 19.5 on Tensor Cores
NVLink on the card2-way bridge, 112.5 GB/s2-GPU bridge, 600 GB/s600 GB/s
Max power300 W300 W400 W
Form factor and coolingDual-slot PCIe card, fanDual-slot PCIe card, passiveSXM4 module

The A6000 figures come from NVIDIA's RTX A6000 datasheet and its GA102 architecture whitepaper, the A100 figures from NVIDIA's A100 datasheet, and compute capability from NVIDIA's CUDA GPU list. NVIDIA also quotes Tensor Core throughput with sparsity, at twice these values. I use the dense figures for all three so the columns compare like with like.

The two A100s share every compute rating. The SXM4 module has 5% more bandwidth and a 400 W power limit instead of 300 W. On 2026-09-27 they cost about the same per GPU, $1.475 to $1.50, so pick by VM size: the SXM4 came in 1 and 8-GPU VMs and the PCIe card in 2 and 4-GPU VMs. The configurations are on the A100 page and the RTX A6000 page.

Prices per GPU-hour#

The A100 costs about three times as much per hour, and the live table shows today's gap.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
A100 SXM4 80GB80 GB$1.50/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes

Prices checked 5 Oct 2026, 04:55 UTC

On 2026-09-27 a 1x A6000 listed at $0.48 an hour, a 1x A100 SXM4 at $1.50 and the A100 PCIe at $1.475 per GPU-hour (today $0.48/GPU-hr, $1.50/GPU-hr and $1.48/GPU-hr). At those prices the A100 SXM4 costs less per job only if it finishes 3.1 times as fast (our calculation: 1.50 / 0.48 = 3.1). On paper it is 2.0 times as fast for Tensor Core math (312 / 154.8) and 2.7 times as fast for work limited by memory bandwidth (2,039 / 768). Neither ratio reaches 3.1. How billing works is on the pricing page, and if you are weighing a purchase instead, the A100 price guide covers buying against renting.

Choose the RTX A6000 when the job fits in 48 GB#

The A6000 wins on cost whenever the job fits, because the A100's speed advantage is smaller than its price premium.

That covers most image generation and a good share of LLM work. SDXL runs on 8 GB cards, FLUX.1 [dev] ships as a 12.3 GB FP8 file, and 7B to 14B models in BF16 or 32B models at 4-bit all load with room to spare. For batches of images or a long QLoRA run on a model up to about 32B, the A6000 should finish the work for less money, even though it finishes later. It also has twice the A100's plain FP32 rate, 38.7 TFLOPS against 19.5, for code that runs on CUDA cores rather than Tensor Cores.

It is also the right card for hours you do not spend computing. A Jupyter notebook, a ComfyUI graph you are still building or a script you are debugging leaves the GPU idle most of the time, and every idle hour on the A6000 cost a third of an A100 hour on 2026-09-27.

When a model outgrows 48 GB, price two A6000s before you rent one A100. A 2-GPU A6000 VM listed at $0.96 per hour on 2026-09-27, with 96 GB across the two cards, against $1.50 for a single A100 SXM4. The trade is that the model is split across two GPUs. Check the link between them with nvidia-smi topo -m after launch, and if it is PCIe rather than NVLink, use pipeline parallelism, which is what vLLM recommends for GPUs without NVLink.

Choose the A100 when the job needs 80 GB or the clock matters#

The A100 wins on memory first. A 70B model at 4-bit is the standard example: the widely used AWQ build of Llama-3.3-70B is a 39.8 GB download. On the 46,068 MiB an A6000 reports with ECC on, vLLM's default of reserving 92% leaves about 4.4 GiB for KV cache, roughly 14,000 tokens shared by every request. On the A100 it leaves about 36.6 GiB, roughly 120,000 tokens (our calculation at 320 KiB of BF16 KV cache per token). Fine-tuning shows the same gap: Unsloth's benchmark gives QLoRA on Llama 3.3 70B a maximum context of 12,106 tokens on a 48 GB GPU and 89,389 on 80 GB. Some models do not load on one A6000 at all, including gpt-oss-120b at about 65 GB and Qwen3-32B in BF16 at 65.5 GB. The KV cache guide explains why context length eats memory this fast.

The second reason is time. NVIDIA describes token-by-token generation as memory-bound, so for a single chat stream bandwidth is the number to compare, and the A100 SXM4 has 2.7 times the A6000's. If a person is waiting on every response, or a training run has a deadline, the A100 buys wall-clock time even though each unit of work costs more.

The third is double precision. The A100 runs FP64 at 9.7 TFLOPS, and at 19.5 on its Tensor Cores. The A6000 runs FP64 at 1/64 of its FP32 rate, and NVIDIA's GA102 whitepaper notes that these cards have no FP64 Tensor Core support. For simulation code that needs FP64, the A6000 is not a real option.

What fits in 48 GB and in 80 GB#

The capacity gap decides more of these choices than speed does. vLLM reserves 92% of GPU memory by default, about 41.4 GiB on the A6000 and 73.6 GiB on the A100 (our calculation from the 46,068 and 81,920 MiB that nvidia-smi reports). The A6000 figure assumes ECC is on, and with ECC off the card reports 49,140 MiB, 2.8 GiB more.

WorkloadMemory it needsSourceRTX A6000, 48 GBA100, 80 GB
SDXL 1.0 imagesRuns on 8 GB cardsStability AIFitsFits
FLUX.1 [dev] images12.3 GB FP8 file, about 23 GB at full precisionBlack Forest Labs, ComfyUI docsFitsFits
Qwen3-8B, BF16, 32k context16.4 GB of weights plus 4.8 GB of KV cacheOur calculation from the model configFitsFits
Llama-3.3-70B, 4-bit AWQ39.8 GB checkpointHugging FaceLoads, about 4.4 GiB left for KV cacheFits, about 36.6 GiB left for KV cache
Qwen3-32B, BF1665.5 GB of weightsOur calculationDoes not fitFits, about 12.6 GiB left for KV cache
gpt-oss-120bAbout 65 GB on diskOpenAIDoes not fitFits on one GPU
Wan2.2 A14B video, 720P59.8 GB peak (official script with offloading)Wan-AIDoes not fitFits
LoRA in 16-bit, 30 to 34B model64 to 80 GBAxolotlDoes not fitTight
QLoRA, 70B model40 to 48 GBAxolotlTightFits

ComfyUI's FP8 files and offloading run the video models in less memory than the official scripts, at the cost of speed. For any model not in the table, the VRAM guide walks through the arithmetic, and the fine-tuning guide covers LoRA and QLoRA.

FAQ#

Is the RTX A6000 a datacenter GPU?

No, it is a workstation card with a fan, a 300 W limit and four DisplayPort outputs. The A100 is built for servers, as a passive PCIe card or an SXM4 module. On QuantaCloud both run the same way: an Ubuntu 22.04 VM with the NVIDIA driver and Docker, reached over SSH as the ubuntu user.

Can either card run FP8 models?

Not natively. FP8 Tensor Core math needs compute capability 8.9, and these are 8.6 and 8.0. vLLM still loads FP8 checkpoints on both as weight-only W8A16: the weights take half the memory of BF16, but the math runs in 16-bit. For native FP8 on a 48 GB card, see L40S vs A100 or RTX 6000 Ada vs RTX A6000.

The cards support it, but check the VM you launch. Run nvidia-smi topo -m: NV followed by a number between two GPUs means NVLink, while PIX, PXB, PHB, NODE or SYS mean the path runs over PCIe. The legend printed under the matrix explains each code.

What happens to my data when I stop the VM?

It is deleted. Stopping an instance terminates it and deletes its disk, and there are no volumes or snapshots, so copy checkpoints and outputs off first. Unused seconds of the current hour are refunded when you stop, as the pricing page explains.


The rule I follow is: start on the A6000, and move to the A100 when 48 GB stops being enough or when someone is waiting on the result. Before you rent either, add up the weights, the KV cache for your context and the batch you plan to run. If the total is under about 44 GB, the A6000 is the better buy. If it is over, or if the job is a deadline rather than a budget, the A100 is. The deploy guide covers launching and connecting.

Launch an RTX A6000 Launch an A100 80GB

Keep building

Choose your next step.