The RTX 6000 Ada is the 48 GB GPU I pick when a model runs in FP8. It pairs the RTX A6000's 48 GB with Ada's FP8 Tensor Cores and 25% more bandwidth, and on QuantaCloud it runs from $0.78/GPU-hr in VMs with 1, 2, 4 or 8 cards. On 2026-09-27 its 8x VM was the lowest-priced 8-GPU configuration in the catalog.
Each configuration is a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu, and its local disk is included in the hourly price. On 2026-09-27 the 1x came with 12 vCPUs, 72 GB of RAM and 350 GB of disk, and the 8x with 104 vCPUs, 640 GB of RAM and 2,800 GB. Offers were in us-midwest-1 and us-midwest-2 that day, with the 8x in us-midwest-1 only. The first hour is charged at launch and the unused seconds of the current hour are refunded when you stop, as the pricing page explains. Every other card is in the GPU catalog.
RTX 6000 Ada specs#
The RTX 6000 Ada Generation is a 300 W Ada workstation card, and its AI numbers are 48 GB, 960 GB/s and FP8.
| Spec | RTX 6000 Ada Generation |
|---|---|
| Architecture | NVIDIA Ada Lovelace, compute capability 8.9 |
| GPU memory | 48 GB GDDR6 with ECC, 384-bit interface |
| Memory bandwidth | 960 GB/s |
| CUDA cores | 18,176 |
| Tensor Cores | 568, fourth generation, with FP8 |
| RT Cores | 142, third generation |
| FP32 | 91.1 TFLOPS |
| Tensor performance | 1,457 TFLOPS FP8 with sparsity, about 728 without (our calculation: half) |
| NVLink | No |
| System interface | PCIe 4.0 x16 |
| Power | 300 W total board power |
| Cooling and size | Active, dual slot, 4.4 x 10.5 in |
| Video engines | 3 encode and 3 decode, with AV1 |
The name trips people up. "RTX 6000 Ada" is not the RTX A6000, which is the Ampere card from the generation before (compute capability 8.6 against 8.9 here), and it is not the RTX PRO 6000 Blackwell, which has 96 GB. The Quadro RTX 6000 is an older Turing card.
What fits on one card, and on eight#
FP8 is what changes the math on this card. An FP8 checkpoint needs half the memory of BF16, and vLLM runs per-channel FP8 checkpoints (the llm-compressor kind) with FP8 math (W8A8) on Ada, while block-scaled ones, such as Qwen's own FP8 releases, run weight-only below Hopper. ComfyUI can compute in FP8 here too, which it cannot do on the A6000. I budget about 44 GB per card for weights and KV cache, the 92% that vLLM claims by default (our calculation: 48 x 0.92 = 44.2 GB). That budget holds with ECC on, when nvidia-smi reports 46,068 MiB, and with ECC off the card reports 49,140 MiB, 2.8 GiB more.
| Workload | Memory it needs | Fits on | Basis |
|---|---|---|---|
| Qwen3-32B in per-channel FP8, one 32k sequence | 34.3 GB checkpoint + 8.6 GB of KV = 42.9 GB | 1x, borderline: just inside the 44.2 GB budget, before activations | Published file size (RedHatAI), our sum |
| Qwen3-32B in BF16, one 32k sequence | About 74 GB | 2x, split across both cards | Our calculation |
| gpt-oss-120b | About 65 GB on disk | 2x, with about 22 GiB left for KV | Published size, our calculation |
| Llama-3.3-70B in FP8 | 72.7 GB checkpoint | 2x with about 48k tokens of BF16 KV, or 4x with a full 128k context (115.6 GB) | Our calculation |
| Llama-3.3-70B in BF16 | 141.1 GB of weights | 4x with about 112k tokens of KV, or 8x for many long contexts | Our calculation |
| FLUX.1 [dev], BFL FP8 file | 12.3 GB | 1x, on a card that can compute in FP8 | Published file size |
| FLUX.2 [dev], fp8 transformer plus fp8 text encoder | 35.5 GB + 18.0 GB = 53.5 GB of files | 1x only with ComfyUI offloading, or 47.7 GB with the 12.3 GB fp4 text encoder, which is tight | Published file sizes, our sum |
| Wan 2.2 A14B video at 480P | 41.3 GB peak on one GPU, with offload | 1x | Published (Wan) |
| QLoRA fine-tuning, 70B | 40 to 48 GB on one GPU | 1x, tight, or 2x with FSDP + QLoRA | Published (Axolotl) |
| Full fine-tuning, Qwen3-8B | 147 GB of model states at 18 bytes per parameter | 8x with ZeRO-3 or FSDP: about 18.4 GB per GPU before activations | Our calculation (8.19B x 18 / 8) |
Splitting a model across cards has a cost on this GPU. The RTX 6000 Ada has no NVLink, so the cards in a 2x, 4x or 8x VM talk over PCIe, and vLLM's docs recommend pipeline parallelism over tensor parallelism when GPUs have no NVLink. Start with --pipeline-parallel-size equal to the GPU count and compare it with --tensor-parallel-size on your own prompts. The vLLM Docker guide, the gpt-oss GPU requirements and the fine-tuning VRAM guide have the details, and the KV cache explainer shows how the context numbers are worked out.
RTX 6000 Ada vs RTX A6000, L40 and L40S#
Among the 48 GB cards, the RTX 6000 Ada has the most bandwidth and ties the L40S on FP8.
| RTX A6000 | RTX 6000 Ada | L40 | L40S | |
|---|---|---|---|---|
| Architecture | Ampere | Ada | Ada | Ada |
| Memory | 48 GB GDDR6 | 48 GB GDDR6 | 48 GB GDDR6 | 48 GB GDDR6 |
| Bandwidth | 768 GB/s | 960 GB/s | 864 GB/s | 864 GB/s |
| FP8 Tensor, dense | None | About 728 TFLOPS (our calculation) | 362 TFLOPS | 733 TFLOPS |
| Max power | 300 W | 300 W | 300 W | 350 W |
| Cooling | Active | Active | Passive | Passive |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
I choose the RTX 6000 Ada over the RTX A6000 whenever the model runs in FP8 or the job leans on bandwidth. Against the L40S, the datasheets tie on FP8 and the RTX 6000 Ada has more bandwidth, so I take whichever has the lower live price. The L40 lists half their FP8 throughput. For the consumer cards people compare it with, see RTX 6000 Ada vs RTX 5090 and RTX 6000 Ada vs RTX 4090.
Where the RTX 6000 Ada falls short#
The RTX 6000 Ada falls short when a model needs one bigger GPU. Each card has 48 GB, there is no NVLink between cards, and there is no FP4, which is Blackwell-only. 8-GPU VMs also take longer to start: across the catalog, the median is about 10 minutes, against about 3 for a single GPU.
| Step up to | Memory and bandwidth | Step up when |
|---|---|---|
| RTX PRO 6000 Blackwell | 96 GB GDDR7, 1,597 or 1,792 GB/s by edition, FP4 | A 70B FP8 model or Wan 2.2 at 720P has to sit on one GPU |
| H100 PCIe | 80 GB HBM2e, 2,000 GB/s | Serving is limited by bandwidth and the model fits in 80 GB |
| H200 NVL | 141 GB HBM3e, 4.8 TB/s | A 70B-class model with long contexts on one GPU |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
The 8x is where this card earns its place. On 2026-09-27 the 8x RTX 6000 Ada cost $6.21 per hour for 384 GB of GPU memory, about 1.6 cents per GB-hour and the lowest of the three 8-GPU configurations in the catalog that day. The 2x H200 NVL cost $6.86 per hour for 282 GB, about 2.4 cents per GB-hour (our calculations: price divided by total GPU memory). The H200 pair gives each GPU five times the bandwidth and leaves fewer GPUs to split a model across, and the 8x gives more memory for the money. For more than eight GPUs or an InfiniBand fabric, we build GPU clusters to order.
Launch the RTX 6000 Ada with a template#
Pick the template by the first job you will run. The template does not change the price, and none of them comes with models preinstalled.
| Template | What you get | Launch |
|---|---|---|
| Bare Metal | Ubuntu 22.04, the NVIDIA driver and Docker over SSH. The start for FP8 serving with vLLM in Docker | Ubuntu + Docker on RTX 6000 Ada |
| PyTorch + Jupyter | JupyterLab with PyTorch in the browser, behind your QuantaCloud login | Jupyter on RTX 6000 Ada |
| Open WebUI + Ollama | A private chat UI. Pull models from inside the app. Open WebUI also has its own sign-in | Open WebUI on RTX 6000 Ada |
| ComfyUI | Image and video generation in the browser, on a card that can compute in FP8. Upload or download your models after launch | ComfyUI on RTX 6000 Ada |
The templates docs list what each template contains, and running ComfyUI on a cloud GPU and the ComfyUI page cover the image and video side.
FAQ#
Is the RTX 6000 Ada the same as the RTX A6000?
No. Both have 48 GB, but the RTX 6000 Ada is the newer Ada card with FP8, 960 GB/s and 18,176 CUDA cores, and the RTX A6000 is Ampere with 768 GB/s, no FP8 and 10,752 CUDA cores. The RTX 6000 Ada vs RTX A6000 comparison asks whether the newer card is worth it.
RTX 6000 Ada or L40S?
On NVIDIA's datasheets they tie on FP8 throughput, and the RTX 6000 Ada has 960 GB/s against the L40S's 864 GB/s. Both have 48 GB, and on 2026-09-27 both came in VMs of up to 8, so I pick on the live price. L40S vs RTX 6000 Ada compares the two in full.
Can I get 8 RTX 6000 Ada GPUs in one VM?
Yes. On 2026-09-27 the catalog offered 1, 2, 4 and 8 per VM, and the 8x came with 104 vCPUs, 640 GB of RAM and 2,800 GB of disk. Expect it to take longer to start than a single GPU, and check the live table at the top for what is available now.
Does the RTX 6000 Ada have NVLink?
No. NVIDIA lists NVLink as "No" for this card, so multi-GPU VMs connect over PCIe.
Is it a VM or bare metal?
A VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu. The template called Bare Metal is the plain Ubuntu image, not bare-metal hardware. Dedicated servers are something we build to order: see dedicated GPU servers.
What happens when I stop?
The VM is terminated and its disk deleted, with no volumes or snapshots to keep data, so copy your outputs off first. The unused seconds of the current hour are refunded.
My rule for the RTX 6000 Ada: rent it when your model runs in FP8, or when the job needs several 48 GB cards in one VM. If the model runs in BF16 or 4-bit on one card, the RTX A6000 does the same work, and if it needs one GPU bigger than 48 GB, go to the RTX PRO 6000 Blackwell. Launch an RTX 6000 Ada , and the vLLM Docker guide gets an FP8 model serving.