On-demand GPU

NVIDIA RTX A6000

48 GB. Room to build.

48 GB of GPU memory for your models, notebooks, and creative workflows. Pick a configuration and make it your own.

Full SSH accessYour choice of template
Starting from$0.48/ GPU-hr
Launch RTX A6000
Available nowSee configurations
NVIDIA RTX A6000AMPERE
VRAMVRAMVRAMVRAMVRAMVRAMAMPERE
Room for the work ahead.48 GB GDDR6
GPU memory
48 GB GDDR6
Memory bandwidth
768 GB/s
Architecture
Ampere
Memory protection
ECC supported

Your GPU. Your configuration.

Choose where you start.

24 configurations
1 GPULowest price

1x RTX A6000

Midwest · us-midwest-1

vCPU
6
RAM
48 GB
Disk
256 GB
$0.48/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-2

vCPU
6
RAM
48 GB
Disk
256 GB
$0.48/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-1

vCPU
6
RAM
24 GB
Disk
256 GB
$0.55/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-2

vCPU
6
RAM
24 GB
Disk
256 GB
$0.55/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-1

vCPU
12
RAM
64 GB
Disk
256 GB
$0.57/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-2

vCPU
12
RAM
64 GB
Disk
256 GB
$0.57/ instance-hour
Available nowLaunch configuration
1 GPU

1x RTX A6000

Midwest · us-midwest-3

vCPU
12
RAM
64 GB
Disk
256 GB
$0.57/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-1

vCPU
14
RAM
96 GB
Disk
512 GB
$0.96/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-2

vCPU
14
RAM
96 GB
Disk
512 GB
$0.96/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-2

vCPU
14
RAM
48 GB
Disk
512 GB
$1.08/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-1

vCPU
14
RAM
48 GB
Disk
512 GB
$1.08/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-1

vCPU
30
RAM
128 GB
Disk
512 GB
$1.13/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-3

vCPU
30
RAM
128 GB
Disk
512 GB
$1.13/ instance-hour
Available nowLaunch configuration
2 GPUs

2x RTX A6000

Midwest · us-midwest-2

vCPU
30
RAM
128 GB
Disk
512 GB
$1.13/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-1

vCPU
30
RAM
192 GB
Disk
1,024 GB
$1.92/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-2

vCPU
30
RAM
192 GB
Disk
1,024 GB
$1.92/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-1

vCPU
30
RAM
96 GB
Disk
1,024 GB
$2.16/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-2

vCPU
30
RAM
96 GB
Disk
1,024 GB
$2.16/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-1

vCPU
54
RAM
256 GB
Disk
1,024 GB
$2.25/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-3

vCPU
54
RAM
256 GB
Disk
1,024 GB
$2.25/ instance-hour
Available nowLaunch configuration
4 GPUs

4x RTX A6000

Midwest · us-midwest-2

vCPU
54
RAM
256 GB
Disk
1,024 GB
$2.25/ instance-hour
Available nowLaunch configuration
8 GPUs

8x RTX A6000

Midwest · us-midwest-1

vCPU
62
RAM
384 GB
Disk
2,560 GB
$3.84/ instance-hour
Available nowLaunch configuration
8 GPUs

8x RTX A6000

Midwest · us-midwest-1

vCPU
62
RAM
192 GB
Disk
1,024 GB
$4.33/ instance-hour
Available nowLaunch configuration
8 GPUs

8x RTX A6000

Midwest · us-midwest-2

vCPU
110
RAM
512 GB
Disk
2,560 GB
$4.50/ instance-hour
Available nowLaunch configuration

Prices checked

The RTX A6000 is the 48 GB GPU I start on. It carries the same 48 GB as the RTX 6000 Ada, the L40 and the L40S, and it was the lowest-priced of the four per GPU-hour when the catalog was checked on 2026-09-27. On QuantaCloud an RTX A6000 cloud VM runs from $0.48/GPU-hr, with 1, 2 or 4 cards per VM. The catch is FP8: this is an Ampere card, so FP8 files save memory on it but do not make the math faster.

Every configuration above is a VM running Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect over SSH as ubuntu. Each one comes with its own vCPUs, RAM and local disk, and the disk is included in the hourly price: on 2026-09-27 the 1x had 6 to 12 vCPUs, 24 to 64 GB of RAM and 256 GB of disk, and the 4x had 30 vCPUs, 192 GB of RAM and 1,024 GB. That day the offers were in our Midwest regions, us-midwest-1 and us-midwest-2. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. The pricing page has the full rules, and the GPU catalog lists every other card.

RTX A6000 specs#

The RTX A6000 is a 300 W Ampere workstation card with 48 GB of ECC memory, and the two figures that matter most for AI work are its 768 GB/s of bandwidth and the FP8 support it does not have.

SpecRTX A6000
ArchitectureNVIDIA Ampere (GA102), compute capability 8.6
GPU memory48 GB GDDR6 with ECC, 384-bit interface
Memory bandwidth768 GB/s
CUDA cores10,752
Tensor Cores336, third generation. No FP8 or FP4
RT Cores84, second generation
FP3238.7 TFLOPS
Tensor performance309.7 TFLOPS, NVIDIA's figure with sparsity
NVLink on the card2-way bridge, 112.5 GB/s bidirectional, bridge sold separately
System interfacePCIe 4.0 x16
Power300 W total board power
Cooling and sizeActive fan, dual slot, 4.4 x 10.5 in

Four NVIDIA names look alike and belong to different generations. The RTX A6000 is Ampere (compute capability 8.6), the RTX 6000 Ada Generation is Ada (8.9), and the RTX PRO 6000 Blackwell is Blackwell (12.0) with 96 GB. The older Quadro RTX 6000 is Turing (7.5).

What fits in 48 GB#

The rule I follow is simple: weights plus KV cache stay under about 44 GB, because vLLM claims 92% of GPU memory by default (our calculation: 48 x 0.92 = 44.2 GB). That budget holds with ECC on, when nvidia-smi reports 46,068 MiB, and with ECC off the card reports 49,140 MiB, 2.8 GiB more. ComfyUI can offload weights to system RAM, so image and video workflows bend that limit at the cost of speed. Each row says where its number comes from.

WorkloadMemory it needsOn one A6000Basis
SDXL 1.0 image generation8 GB of VRAMYesPublished (Stability AI)
FLUX.1 [dev], 12B24 GB of weights in BF16, 12 GB in FP8, before text encodersYes, in either precisionOur calculation (12B x 2 or 1 bytes)
Qwen-Image with its text encoder, fp8 files20.4 GB + 9.4 GB = 29.8 GB of filesYesPublished file sizes, our sum
Wan 2.2 A14B video, 480P and 720P41.3 GB and 59.8 GB peak on one GPU, with offload480P yes, 720P not at Wan's published peakPublished (Wan)
Qwen3-8B in BF16, 32k context16.4 GB of weights plus 4.5 GiB of KV per 32k sequenceYes, about five 32k sequences at onceOur calculation
gpt-oss-20bWithin 16 GB of memoryYesPublished (OpenAI)
Llama-3.3-70B, 4-bit AWQ39.8 GB checkpoint, leaving about 4.4 GiB for KVYes, about 14k tokens of BF16 KVOur calculation
gpt-oss-120bAbout 65 GB on diskNo. It needs 2x A6000Published size, our calculation
QLoRA fine-tuning, 7B to 8B10 to 14 GBYesPublished (Axolotl)
QLoRA fine-tuning, 30B to 34B24 to 32 GBYesPublished (Axolotl)
QLoRA fine-tuning, 70B40 to 48 GB. Unsloth's benchmark reached a 12,106-token context for Llama 3.3 70B at 48 GBTightPublished (Axolotl, Unsloth)

The one thing I always check on this card is the precision of the files. FP8 checkpoints and fp8 ComfyUI files load on the A6000 and halve the weight memory, but ComfyUI only computes in FP8 on compute capability 8.9 and newer, and vLLM runs FP8 checkpoints weight-only (W8A16) on Ampere. On the A6000, FP8 is a memory saving, not a speed-up. NVFP4 files behave the same way: vLLM falls back to weight-only 4-bit kernels. To size a model that is not in the table, the VRAM guide and the KV cache explainer show the arithmetic, and the LoRA, QLoRA and full fine-tuning guide covers training.

Where the A6000 falls short#

The A6000 falls short on FP8 and on bandwidth. NVIDIA describes LLM token generation as memory-bound: the time spent moving weights and KV cache out of memory dominates, not the math. At 768 GB/s the A6000 has the least bandwidth of the 48 GB cards in the catalog. The RTX 6000 Ada has 25% more and the L40S 12.5% more (our calculation from NVIDIA's figures). These are the cards I step up to, and why:

Step up toMemoryBandwidthFP8 and FP4Step up when
RTX 6000 Ada48 GB GDDR6960 GB/sFP8You serve per-channel FP8 checkpoints at the same 48 GB, or want 8 GPUs in one VM (offered on 2026-09-27)
L40S48 GB GDDR6864 GB/sFP8The RTX 6000 Ada is not available: the L40S lists 733 dense FP8 TFLOPS at the same 48 GB
RTX PRO 6000 Blackwell96 GB GDDR71,597 or 1,792 GB/s, by editionFP8 and FP4The model or video workflow does not fit in 48 GB
H200 NVL141 GB HBM3e4.8 TB/sFP8You serve a 70B-class model or long contexts on one GPU
GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

For models between 48 and 80 GB, the A100 80GB is the other route. It has no FP8 either, but it has 80 GB of HBM2e at 1,935 to 2,039 GB/s. The RTX A6000 vs A100 comparison covers that choice, and RTX A6000 vs RTX 5090 covers the consumer card people often weigh it against. RTX A6000 vs RTX 4090 does the same for the RTX 4090.

Measured on QuantaCloud#

Across the catalog, most single-GPU VMs are running in about 3 minutes (median).

Launch the A6000 with a template#

Pick the template by the first job you will run on the machine. The template does not change the price, and none of them comes with models preinstalled.

TemplateWhat you getLaunch
Bare MetalUbuntu 22.04, the NVIDIA driver and Docker over SSH. The start for vLLM in Docker, your own containers and scriptsUbuntu + Docker on A6000
PyTorch + JupyterJupyterLab with PyTorch in the browser, behind your QuantaCloud loginJupyter on A6000
Open WebUI + OllamaA private chat UI. Pull models from inside the app. Open WebUI also has its own sign-inOpen WebUI on A6000
ComfyUINode-based image and video generation in the browser. Upload or download your checkpoints after launchComfyUI on A6000

Bare Metal is the name of the plain Ubuntu template: the instance is still a VM. The templates docs list what each one contains, and deploying a GPU walks through the console. The ComfyUI, Open WebUI and Jupyter pages go further on the app templates, and the Open WebUI and Ollama setup guide walks through the first chat.

FAQ#

Is the RTX A6000 the same as the RTX 6000 Ada?

No. The RTX A6000 is an Ampere card with 768 GB/s and no FP8. The RTX 6000 Ada Generation has the same 48 GB with 960 GB/s, FP8 Tensor Cores and 18,176 CUDA cores to the A6000's 10,752. If you serve per-channel FP8 checkpoints, the RTX 6000 Ada is the better card: vLLM runs them with FP8 math on its Tensor Cores, while block-scaled FP8 checkpoints, such as Qwen's own FP8 releases, run weight-only on both cards. If your models run in BF16 or 4-bit and fit, the A6000 does the same work, and the live price table above shows what each costs today. The RTX 6000 Ada vs RTX A6000 comparison weighs the two side by side.

RTX A6000 or A100 for AI work?

The A100 80GB, when the model needs between 48 and 80 GB or when serving is limited by bandwidth: it has 80 GB of HBM2e at 1,935 to 2,039 GB/s. Neither card has FP8. For anything that fits in 48 GB, I pick the A6000.

Can I get 2 or 4 A6000s in one VM?

Yes. On 2026-09-27 the catalog offered 1, 2 or 4 per VM, and the table at the top shows what is available now. vLLM can split one model across the cards with tensor or pipeline parallelism, and fine-tuning frameworks use every GPU in the VM through DDP or FSDP.

QuantaCloud does not claim it. The A6000 supports a 2-way NVLink bridge, but whether a VM has one shows only in nvidia-smi topo -m.

Without NVLink, vLLM's docs recommend pipeline parallelism over tensor parallelism.

Is it a VM or bare metal?

A VM. Every QuantaCloud instance runs Ubuntu 22.04 with the NVIDIA driver and Docker, and "Bare Metal" is only the name of the plain Ubuntu template. If you need a physical server of your own, we build dedicated GPU servers to order.

What happens to my files when I stop?

They are deleted. Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy results off before you stop. The unused seconds of the current hour are refunded. If your balance cannot cover the next hour, the VM is terminated the same way, and a low-balance email goes out when the balance drops below $2.

How much does an RTX A6000 cost per hour?

From $0.48/GPU-hr. You pay for the time you use: a session stopped after 3 hours and 20 minutes is charged four hours as it runs, then the unused 40 minutes are refunded. At the lowest 1x price on 2026-09-27, $0.48 per hour, that session costs $1.60 (our calculation: 3.33 hours x $0.48).

Is the A6000 available right now?

The configurations table at the top comes from the live catalog, with the time it was checked, and availability changes by the minute. If no A6000 is listed, the GPU catalog shows what is available instead.


My rule for the A6000 is short: if the model and its context fit in about 44 GB and you do not need FP8 speed, rent the A6000 and spend the difference on more hours. If you need FP8 math, take the RTX 6000 Ada, and if you need more than 48 GB on one GPU, go to the RTX PRO 6000 Blackwell. Launch an RTX A6000 , and how to rent a GPU covers the first login.

Keep building

Choose your next step.