The RTX A6000 is the 48 GB GPU I start on. It carries the same 48 GB as the RTX 6000 Ada, the L40 and the L40S, and it was the lowest-priced of the four per GPU-hour when the catalog was checked on 2026-09-27. On QuantaCloud an RTX A6000 cloud VM runs from $0.48/GPU-hr, with 1, 2 or 4 cards per VM. The catch is FP8: this is an Ampere card, so FP8 files save memory on it but do not make the math faster.
Every configuration above is a VM running Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect over SSH as ubuntu. Each one comes with its own vCPUs, RAM and local disk, and the disk is included in the hourly price: on 2026-09-27 the 1x had 6 to 12 vCPUs, 24 to 64 GB of RAM and 256 GB of disk, and the 4x had 30 vCPUs, 192 GB of RAM and 1,024 GB. That day the offers were in our Midwest regions, us-midwest-1 and us-midwest-2. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. The pricing page has the full rules, and the GPU catalog lists every other card.
RTX A6000 specs#
The RTX A6000 is a 300 W Ampere workstation card with 48 GB of ECC memory, and the two figures that matter most for AI work are its 768 GB/s of bandwidth and the FP8 support it does not have.
| Spec | RTX A6000 |
|---|---|
| Architecture | NVIDIA Ampere (GA102), compute capability 8.6 |
| GPU memory | 48 GB GDDR6 with ECC, 384-bit interface |
| Memory bandwidth | 768 GB/s |
| CUDA cores | 10,752 |
| Tensor Cores | 336, third generation. No FP8 or FP4 |
| RT Cores | 84, second generation |
| FP32 | 38.7 TFLOPS |
| Tensor performance | 309.7 TFLOPS, NVIDIA's figure with sparsity |
| NVLink on the card | 2-way bridge, 112.5 GB/s bidirectional, bridge sold separately |
| System interface | PCIe 4.0 x16 |
| Power | 300 W total board power |
| Cooling and size | Active fan, dual slot, 4.4 x 10.5 in |
Four NVIDIA names look alike and belong to different generations. The RTX A6000 is Ampere (compute capability 8.6), the RTX 6000 Ada Generation is Ada (8.9), and the RTX PRO 6000 Blackwell is Blackwell (12.0) with 96 GB. The older Quadro RTX 6000 is Turing (7.5).
What fits in 48 GB#
The rule I follow is simple: weights plus KV cache stay under about 44 GB, because vLLM claims 92% of GPU memory by default (our calculation: 48 x 0.92 = 44.2 GB). That budget holds with ECC on, when nvidia-smi reports 46,068 MiB, and with ECC off the card reports 49,140 MiB, 2.8 GiB more. ComfyUI can offload weights to system RAM, so image and video workflows bend that limit at the cost of speed. Each row says where its number comes from.
| Workload | Memory it needs | On one A6000 | Basis |
|---|---|---|---|
| SDXL 1.0 image generation | 8 GB of VRAM | Yes | Published (Stability AI) |
| FLUX.1 [dev], 12B | 24 GB of weights in BF16, 12 GB in FP8, before text encoders | Yes, in either precision | Our calculation (12B x 2 or 1 bytes) |
| Qwen-Image with its text encoder, fp8 files | 20.4 GB + 9.4 GB = 29.8 GB of files | Yes | Published file sizes, our sum |
| Wan 2.2 A14B video, 480P and 720P | 41.3 GB and 59.8 GB peak on one GPU, with offload | 480P yes, 720P not at Wan's published peak | Published (Wan) |
| Qwen3-8B in BF16, 32k context | 16.4 GB of weights plus 4.5 GiB of KV per 32k sequence | Yes, about five 32k sequences at once | Our calculation |
| gpt-oss-20b | Within 16 GB of memory | Yes | Published (OpenAI) |
| Llama-3.3-70B, 4-bit AWQ | 39.8 GB checkpoint, leaving about 4.4 GiB for KV | Yes, about 14k tokens of BF16 KV | Our calculation |
| gpt-oss-120b | About 65 GB on disk | No. It needs 2x A6000 | Published size, our calculation |
| QLoRA fine-tuning, 7B to 8B | 10 to 14 GB | Yes | Published (Axolotl) |
| QLoRA fine-tuning, 30B to 34B | 24 to 32 GB | Yes | Published (Axolotl) |
| QLoRA fine-tuning, 70B | 40 to 48 GB. Unsloth's benchmark reached a 12,106-token context for Llama 3.3 70B at 48 GB | Tight | Published (Axolotl, Unsloth) |
The one thing I always check on this card is the precision of the files. FP8 checkpoints and fp8 ComfyUI files load on the A6000 and halve the weight memory, but ComfyUI only computes in FP8 on compute capability 8.9 and newer, and vLLM runs FP8 checkpoints weight-only (W8A16) on Ampere. On the A6000, FP8 is a memory saving, not a speed-up. NVFP4 files behave the same way: vLLM falls back to weight-only 4-bit kernels. To size a model that is not in the table, the VRAM guide and the KV cache explainer show the arithmetic, and the LoRA, QLoRA and full fine-tuning guide covers training.
Where the A6000 falls short#
The A6000 falls short on FP8 and on bandwidth. NVIDIA describes LLM token generation as memory-bound: the time spent moving weights and KV cache out of memory dominates, not the math. At 768 GB/s the A6000 has the least bandwidth of the 48 GB cards in the catalog. The RTX 6000 Ada has 25% more and the L40S 12.5% more (our calculation from NVIDIA's figures). These are the cards I step up to, and why:
| Step up to | Memory | Bandwidth | FP8 and FP4 | Step up when |
|---|---|---|---|---|
| RTX 6000 Ada | 48 GB GDDR6 | 960 GB/s | FP8 | You serve per-channel FP8 checkpoints at the same 48 GB, or want 8 GPUs in one VM (offered on 2026-09-27) |
| L40S | 48 GB GDDR6 | 864 GB/s | FP8 | The RTX 6000 Ada is not available: the L40S lists 733 dense FP8 TFLOPS at the same 48 GB |
| RTX PRO 6000 Blackwell | 96 GB GDDR7 | 1,597 or 1,792 GB/s, by edition | FP8 and FP4 | The model or video workflow does not fit in 48 GB |
| H200 NVL | 141 GB HBM3e | 4.8 TB/s | FP8 | You serve a 70B-class model or long contexts on one GPU |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
For models between 48 and 80 GB, the A100 80GB is the other route. It has no FP8 either, but it has 80 GB of HBM2e at 1,935 to 2,039 GB/s. The RTX A6000 vs A100 comparison covers that choice, and RTX A6000 vs RTX 5090 covers the consumer card people often weigh it against. RTX A6000 vs RTX 4090 does the same for the RTX 4090.
Measured on QuantaCloud#
Across the catalog, most single-GPU VMs are running in about 3 minutes (median).
Launch the A6000 with a template#
Pick the template by the first job you will run on the machine. The template does not change the price, and none of them comes with models preinstalled.
| Template | What you get | Launch |
|---|---|---|
| Bare Metal | Ubuntu 22.04, the NVIDIA driver and Docker over SSH. The start for vLLM in Docker, your own containers and scripts | Ubuntu + Docker on A6000 |
| PyTorch + Jupyter | JupyterLab with PyTorch in the browser, behind your QuantaCloud login | Jupyter on A6000 |
| Open WebUI + Ollama | A private chat UI. Pull models from inside the app. Open WebUI also has its own sign-in | Open WebUI on A6000 |
| ComfyUI | Node-based image and video generation in the browser. Upload or download your checkpoints after launch | ComfyUI on A6000 |
Bare Metal is the name of the plain Ubuntu template: the instance is still a VM. The templates docs list what each one contains, and deploying a GPU walks through the console. The ComfyUI, Open WebUI and Jupyter pages go further on the app templates, and the Open WebUI and Ollama setup guide walks through the first chat.
FAQ#
Is the RTX A6000 the same as the RTX 6000 Ada?
No. The RTX A6000 is an Ampere card with 768 GB/s and no FP8. The RTX 6000 Ada Generation has the same 48 GB with 960 GB/s, FP8 Tensor Cores and 18,176 CUDA cores to the A6000's 10,752. If you serve per-channel FP8 checkpoints, the RTX 6000 Ada is the better card: vLLM runs them with FP8 math on its Tensor Cores, while block-scaled FP8 checkpoints, such as Qwen's own FP8 releases, run weight-only on both cards. If your models run in BF16 or 4-bit and fit, the A6000 does the same work, and the live price table above shows what each costs today. The RTX 6000 Ada vs RTX A6000 comparison weighs the two side by side.
RTX A6000 or A100 for AI work?
The A100 80GB, when the model needs between 48 and 80 GB or when serving is limited by bandwidth: it has 80 GB of HBM2e at 1,935 to 2,039 GB/s. Neither card has FP8. For anything that fits in 48 GB, I pick the A6000.
Can I get 2 or 4 A6000s in one VM?
Yes. On 2026-09-27 the catalog offered 1, 2 or 4 per VM, and the table at the top shows what is available now. vLLM can split one model across the cards with tensor or pipeline parallelism, and fine-tuning frameworks use every GPU in the VM through DDP or FSDP.
Is NVLink enabled on the 2x and 4x VMs?
QuantaCloud does not claim it. The A6000 supports a 2-way NVLink bridge, but whether a VM has one shows only in nvidia-smi topo -m.
Without NVLink, vLLM's docs recommend pipeline parallelism over tensor parallelism.
Is it a VM or bare metal?
A VM. Every QuantaCloud instance runs Ubuntu 22.04 with the NVIDIA driver and Docker, and "Bare Metal" is only the name of the plain Ubuntu template. If you need a physical server of your own, we build dedicated GPU servers to order.
What happens to my files when I stop?
They are deleted. Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy results off before you stop. The unused seconds of the current hour are refunded. If your balance cannot cover the next hour, the VM is terminated the same way, and a low-balance email goes out when the balance drops below $2.
How much does an RTX A6000 cost per hour?
From $0.48/GPU-hr. You pay for the time you use: a session stopped after 3 hours and 20 minutes is charged four hours as it runs, then the unused 40 minutes are refunded. At the lowest 1x price on 2026-09-27, $0.48 per hour, that session costs $1.60 (our calculation: 3.33 hours x $0.48).
Is the A6000 available right now?
The configurations table at the top comes from the live catalog, with the time it was checked, and availability changes by the minute. If no A6000 is listed, the GPU catalog shows what is available instead.
My rule for the A6000 is short: if the model and its context fit in about 44 GB and you do not need FP8 speed, rent the A6000 and spend the difference on more hours. If you need FP8 math, take the RTX 6000 Ada, and if you need more than 48 GB on one GPU, go to the RTX PRO 6000 Blackwell. Launch an RTX A6000 , and how to rent a GPU covers the first login.