On-demand NVIDIA GPUs

Find the GPU for your next workload.

Rent NVIDIA GPU VMs by the hour: RTX A6000 to H200 NVL, 48 to 141 GB per GPU, 1 to 8 GPUs per VM, live prices and US regions.

Live prices and availabilityGPU memory, region, and configuration

Your GPU. Your workload.

Choose your GPU.

7 GPU modelsUSD / hour
NVIDIA

RTX A6000

48 GB VRAM per GPU

GPUs per VM
1 / 2 / 4 / 8
Configurations
24
Regions
3

Single configuration from

$0.48/ instance-hour
Available nowSee configurations
NVIDIA

RTX 6000 Ada

48 GB VRAM per GPU

GPUs per VM
1 / 2 / 4 / 8
Configurations
7
Regions
2

Single configuration from

$0.79/ instance-hour
Available nowSee configurations
NVIDIA

L40

48 GB VRAM per GPU

GPUs per VM
1 / 2 / 4
Configurations
3
Regions
1

Single configuration from

$0.94/ instance-hour
Available nowSee configurations
NVIDIA

L40S

48 GB VRAM per GPU

GPUs per VM
1 / 2
Configurations
4
Regions
2

Single configuration from

$1.09/ instance-hour
Available nowSee configurations
NVIDIA

A100

80 GB VRAM per GPU

GPUs per VM
1 / 2 / 8
Configurations
7
Regions
3

Single configuration from

$1.48/ instance-hour
Available nowSee configurations
NVIDIA

RTX PRO 6000 Blackwell

96 GB VRAM per GPU

GPUs per VM
1
Configurations
2
Regions
2

Single configuration from

$2.39/ instance-hour
Available nowSee configurations
NVIDIA

H100 PCIe

80 GB VRAM per GPU

GPUs per VM
1
Configurations
1
Regions
1

Single configuration from

$2.59/ instance-hour
Available nowSee configurations

Prices checked

Plan what comes next

Dedicated hardware, built to order.

Plan your capacity

Rent an NVIDIA GPU by the hour on QuantaCloud: pick a card with 48 to 141 GB of memory, launch an Ubuntu VM or a ready-made template, and pay for the time between launch and stop. Every instance is a virtual machine provided by QuantaCloud in our US regions, Virginia and the Midwest, with 1, 2, 4 or 8 GPUs. The catalog below comes from the console's live offers, stamped with the time it was checked.

Launch an RTX A6000

How billing works

GPUs available now#

The live table answers what you can rent today: each row is a GPU family with its memory, the lowest current price per GPU-hour, the GPU counts on offer and the regions it runs in.

Availability moves with capacity, and a GPU can show as unavailable while its offers are taken. If the GPU you need is missing, or you need it for months rather than hours, send a capacity brief and we will quote a reserved build in writing.

The GPUs side by side#

Memory decides most choices, and bandwidth and FP8 support decide the rest. The figures are NVIDIA's.

GPUArchitectureMemory per GPUMemory bandwidthFP8 Tensor Cores
RTX A6000Ampere48 GB GDDR6768 GB/sNo
RTX 6000 AdaAda Lovelace48 GB GDDR6960 GB/sYes
L40Ada Lovelace48 GB GDDR6864 GB/sYes
L40SAda Lovelace48 GB GDDR6864 GB/sYes
A100 80GB, SXM4 or PCIeAmpere80 GB HBM2e2,039 GB/s (SXM4), 1,935 GB/s (PCIe)No
H100 PCIeHopper80 GB HBM2e2,000 GB/sYes
RTX PRO 6000 BlackwellBlackwell96 GB GDDR71,597 or 1,792 GB/s, by editionYes, plus FP4
H200 NVLHopper141 GB HBM3e4,813 GB/sYes

The FP8 column matters when your model ships in FP8: FP8 Tensor Core math needs compute capability 8.9 or higher, and the Ampere cards sit below it (8.6 for the RTX A6000, 8.0 for the A100). On the Ada cards, vLLM uses FP8 math for per-channel FP8 checkpoints, while block-scaled ones, such as Qwen's own FP8 releases, run weight-only below Hopper. For BF16 and FP16 work that fits in 48 GB, the RTX A6000 is the card I would start on: it was the lowest-priced GPU in the catalog on 2026-09-27.

Which GPU for which job#

The rule I follow: pick the smallest memory tier that holds the model, its context and some margin, and add GPUs only when one card is not enough. The figures below are published by the model makers or tool authors, or are our calculations from their configs.

Memory per GPUGPUsWhat fits on one GPU
48 GBRTX A6000, RTX 6000 Ada, L40, L40SSDXL, which Stability says runs on 8 GB consumer cards. FLUX.1 [dev] at full 16-bit precision, 34.2 GB of model, text-encoder and VAE files (our sum of the Hugging Face file sizes). Qwen-Image in FP8, which ComfyUI's docs ran on a 24 GB card. Llama 3.1 8B in BF16, 16 GB of weights. QLoRA on a 70B model, 40 to 48 GB by Axolotl's estimate, which leaves no margin.
80 GBA100 80GB, H100 PCIegpt-oss-120b, which OpenAI says fits a single 80 GB GPU. Wan 2.2 A14B video at 720P, which Wan says can run on a GPU with at least 80 GB, with offloading. QLoRA on a 70B model, with room to spare.
96 GBRTX PRO 6000 BlackwellQwen3-32B in BF16 with one 32,768-token sequence, about 74 GB before activations (our calculation: 65.5 GB of weights plus 8.6 GB of KV cache). Tight on an 80 GB card, comfortable here.
141 GBH200 NVLLlama 3.3 70B in FP8 with one full 128k-token context, about 115.6 GB (our calculation: a 72.7 GB FP8 checkpoint plus 42.9 GB of BF16 KV cache).

A model that does not fit on one card can be split across the GPUs of one VM. On cards without NVLink, such as the L40S, vLLM's documentation recommends pipeline parallelism over tensor parallelism. For a model that is not in the table, how much VRAM you need walks through the arithmetic, and the KV cache guide covers the part that grows with context length.

What every instance includes#

Every instance is an Ubuntu 22.04 VM with the NVIDIA driver and Docker installed.

PartWhat you get
AccessSSH as ubuntu on port 22, with your key, at a public IP address
DiskA fixed local disk included in the price. On 2026-09-27 it ran from 256 GB with one RTX A6000 to 5,000 GB with eight A100 SXM4 or L40S GPUs
NetworkNo egress or ingress charges
TemplatesComfyUI, Open WebUI + Ollama, PyTorch + Jupyter, or Bare Metal (plain Ubuntu with the driver and Docker). The template does not change the price
Browser appsApp templates open at a private address behind your QuantaCloud login, and only your account can open them

Bare Metal is the name of the plain template, not the hardware: every on-demand instance is a VM. The ComfyUI, Open WebUI and Jupyter pages cover the app templates, the GPU VPS page covers the plain VM, and the templates docs list what each one runs.

How billing works#

You pay for the seconds between launch and stop, prepaid one hour at a time. You add credit by card first, $5 at minimum. Launching charges the first hour, each further hour is charged when the previous one is used up, and when you stop, the unused seconds come back to your balance. If you stop before the instance is running, or provisioning fails, the whole first hour is refunded.

On 2026-09-27 a single RTX A6000 cost $0.48 an hour (today: $0.48/GPU-hr). A 90-minute session at that price is charged $0.96 over two hours and refunded $0.24 on stop, so it costs $0.72 (our calculation: 1.5 x $0.48). The pricing page has the full rules and three worked examples.

Good to know before you launch#

Stopping an instance deletes it, disk included. There are no volumes and no snapshots, so copy your results off before you press Stop. The same happens when your balance cannot cover the next hour: the instance is terminated and its disk is gone, so I would turn on auto top-up for any run you are not watching.

QuantaCloud runs in the US only: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 in the Midwest. There is no SLA, and support is by email and an in-console ticket.

Some hardware is not on demand at all: 8x H100 or H200, B200, B300 and multi-node GPU clusters. We build those to order as reserved capacity, with the configuration, lead time and terms confirmed in writing. Send a capacity brief.

Launch in three steps#

  1. Create an account in the QuantaCloud console with GitHub, Google, an emailed magic link, or an email and password.
  2. Add credit on the Billing page, $5 at minimum (adding credits).
  3. Pick an offer in the marketplace, choose a template and deploy (deploying a GPU).

The walkthrough in how to rent a GPU for AI covers each step, including the SSH key to save on the first day.

Frequently asked questions#

Is this a VM or bare metal?

A VM. Every on-demand instance is a virtual machine running Ubuntu 22.04 with the NVIDIA driver and Docker. Physical servers are available only as reserved builds, described on the dedicated GPU servers page.

What happens when I stop an instance?

It is terminated and its disk is deleted. The unused seconds of the current hour go back to your balance. To work again, launch a new instance, which starts clean.

Can I get 8 GPUs in one VM?

Yes, on some GPUs. On 2026-09-27 the 8-GPU VMs were the A100 SXM4 80GB, the L40S and the RTX 6000 Ada, and the live catalog above shows today's counts. Anything larger than one 8-GPU VM is a reserved cluster.

Do you have H100 SXM, B200 or B300 on demand?

No. We build them to order as reserved capacity, from a single HGX server to an InfiniBand cluster. The B200 and B300 pages have the specs, and the reserved capacity page has the brief form.

Which regions can I choose?

Four US regions: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 in the Midwest. The region is part of each offer, so you pick it when you pick the offer.

Can I reserve capacity instead of paying by the hour?

Yes. Reserved capacity is built to order: any NVIDIA GPU, from one server to an InfiniBand cluster, quoted in writing. Send a capacity brief.


My decision rule: start on one RTX A6000 when the job fits in 48 GB, take an L40S or RTX 6000 Ada instead when it needs FP8 math, step up to 80 GB for gpt-oss-120b or a 70B fine-tune with headroom, and go to the H200 NVL only when one model needs more than 96 GB.

Launch an RTX A6000

Keep building

Choose your next step.