GPU as a service (GPUaaS) is renting GPU computing over the network instead of buying the hardware. You launch GPU servers or virtual machines when you need them, pay by the hour or for an agreed term, and hand them back when you are done, so the hardware, its power and its cooling stay off your books. It is the GPU form of infrastructure as a service, and it is sold in four models: on-demand, reserved, dedicated and spot.
What GPU as a service means#
The clearest definition predates the GPU boom: NIST's definition of cloud computing, published in 2011. It describes infrastructure as a service as "the capability provided to the consumer is to provision processing, storage, networks, and other fundamental computing resources where the consumer is able to deploy and run arbitrary software". GPUaaS is that capability with GPUs as the processing. You bring the software, the models and the data, and the provider brings the machine.
Two of NIST's five essential characteristics matter most when you compare GPU offers. On-demand self-service means a consumer "can unilaterally provision computing capabilities, such as server time and network storage, as needed automatically without requiring human interaction with each service provider", and measured service means the provider meters what you use, typically to charge you per use. On QuantaCloud, the first is a console where you pick a GPU offer and click Deploy, and the second is a bill that runs from launch to stop. The NIST definition is short and worth reading once.
The four ways GPU compute is sold#
GPU compute is sold in four models, and they differ in commitment far more than in hardware.
| Model | What you get | How it is priced | Commitment | Fits |
|---|---|---|---|---|
| On-demand | A VM or server you launch and stop yourself | Per GPU-hour, often metered more finely | None | Experiments, bursts and work measured in hours or weeks |
| Reserved | Capacity held for you for a term | A quote for the configuration and the term | A contract term | Steady work you can forecast |
| Dedicated | Single-tenant physical servers, often bare metal | A quote per server and term | A contract term | Full control of the hardware, long training runs, compliance needs |
| Spot or interruptible | Spare capacity the provider can take back | A lower hourly price | None, but the instance can be reclaimed | Jobs that checkpoint often and can restart |
A fifth product often goes by the same name: model APIs priced per token, where the provider runs the model and you never see the GPU. That is a model service rather than GPU compute. It saves you the setup, and in exchange the provider picks the models and your prompts travel to its servers.
QuantaCloud sells the first three. On-demand instances are VMs with 1, 2, 4 or 8 NVIDIA GPUs, from the 48 GB RTX A6000 to the 141 GB H200 NVL, in US regions, and the GPU catalog shows what is available with live prices. Reserved capacity and dedicated GPU servers are built to order: any NVIDIA GPU configuration, from one server to a GPU cluster with InfiniBand, with the configuration, lead time and terms in writing before you commit. There are no spot instances and no per-token model APIs.
How GPUaaS pricing works#
The unit is the GPU-hour: an hourly price for each GPU, times the number of GPUs, times the hours. Two offers at the same GPU-hour price can still cost very different amounts, because of what sits around that price.
| What to check | Why it changes the bill | On QuantaCloud |
|---|---|---|
| Billing granularity and minimums | A 10-minute test billed as a full hour pays for 6 times the time it used (our calculation: 60 / 10) | The first hour is charged at launch, and the unused seconds are refunded when you stop |
| What the hourly price includes | CPU, RAM and disk can be billed separately | vCPU, RAM and a local disk come with each configuration, for example 16 vCPU, 180 GB of RAM and 750 GB of disk with one H200 NVL on 2026-09-27 |
| Data transfer | Moving datasets and checkpoints out can be charged per GB | No egress or ingress charges |
| Storage after you stop | Volumes that outlive an instance are billed per GB per month | No volumes or snapshots: stopping deletes the disk |
| Payment | Prepaid credit or a monthly invoice | Prepaid credit by card, from $5, with optional auto top-up |
| Commitment | A term buys capacity held for you | None on demand. Reserved terms are quoted in writing |
Today's lowest price per GPU-hour for each GPU family on QuantaCloud:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| H100 PCIe | - | Not listed | No |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 5 Oct 2026, 20:25 UTC
The pricing page has QuantaCloud's billing rules in full, including what happens when a balance runs out: the instance is terminated at the first hourly charge the balance cannot cover, and its disk is deleted.
On-demand, reserved or buy: utilization decides#
Utilization decides it: the share of hours a GPU spends on useful work. An always-on month is 730 hours (8,760 hours a year divided by 12), so at the 2026-09-27 prices one RTX A6000 on demand costs $350.40 a month and one H200 NVL $2,503.90 (our calculation: 730 x $0.48 and 730 x $3.43). The same RTX A6000 used 8 hours a day on 22 working days costs $84.48 (our calculation: 176 x $0.48). Today's prices are $0.48/GPU-hr and the console price. Part-time use is where on-demand wins, as long as your setup is a script, because each QuantaCloud session starts from a clean VM.
Buying wins once the hardware stays busy, and how busy depends on what you compare it with. A February 2025 industry analysis of 8-GPU H100 systems found that owning one in a US colocation facility was cheaper than renting on demand from the large hyperscalers once it was busy 22% of the time or more. Against GPU-focused clouds the break-even rose to 66%, because their on-demand price averaged $34 an hour against $98, about a third (our calculation: $34 / $98 = 0.35).
The rule I follow: stay on demand while the load is bursty or the plan changes month to month, ask for a reserved quote once a workload runs most hours of most days for months, and buy hardware only when it will stay busy and you want to run it yourself. The H100 price guide, A100 price guide and H200 price guide work through buying against renting for each GPU.
Which GPU to rent#
Memory decides most choices: the model, its context and the batch have to fit in the GPU's memory, and speed only matters among the GPUs that hold them. As a rough map of QuantaCloud's catalog, a 48 GB card holds an 8B model with long contexts or a 70B model in 4-bit with short ones, an 80 to 96 GB card holds gpt-oss-120b, and the 141 GB H200 NVL holds a 70B model in FP8 with its full 128k context. How much VRAM you need does the arithmetic for any workload, and the LLM inference and fine-tuning pages size the two most common jobs.
Launch an RTX A6000 on demandWhat to compare between GPU clouds#
Compare GPU clouds on the points below, not on the headline price alone. The right column shows how QuantaCloud answers each one, including the answers that are a no.
| What to ask about | Why it matters | QuantaCloud's answer |
|---|---|---|
| GPUs you can launch today | A listing is not the same as capacity you can start now | RTX A6000 to H200 NVL on demand, 48 to 141 GB per GPU, shown live. B200, B300 and 8-GPU H100 or H200 servers are built to order |
| Billing | Granularity and prepayment change the real cost | Prepaid credit, first hour charged at launch, unused seconds refunded on stop |
| Your data when an instance stops | Checkpoints, models and datasets | Stop terminates the instance and deletes its disk. There are no volumes or snapshots |
| How the GPUs connect | Multi-GPU and multi-node jobs | VMs with up to 8 GPUs. Check NVLink with nvidia-smi topo -m. InfiniBand clusters are built to order |
| Where it runs | Latency and data residency | US only: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 |
| Access | Your own drivers, containers and tools | Ubuntu 22.04 VMs with SSH as ubuntu and Docker. Bare-metal servers only as reserved builds |
| Time to start | Time to the first result | Most single-GPU VMs are running in about 3 minutes (median) |
| Service commitment | Production dependency | No SLA. Support is by email and in-console ticket |
| Certifications | Procurement and compliance | Not SOC 2 or HIPAA certified |
How to rent a GPU walks through QuantaCloud's side of this from sign-up to a running instance, and InfiniBand vs NVLink explains the interconnect row.
GPUaaS questions#
Is GPU as a service the same as a GPU cloud?
Close enough. A GPU cloud is the provider, and GPU as a service is what it sells. GPU-focused providers are also called neoclouds, and what is a neocloud explains how they differ from hyperscalers.
Is GPU as a service cheaper than buying GPUs?
Below the break-even utilization, yes. The February 2025 analysis put it at 22% against hyperscaler on-demand prices and 66% against GPU-focused clouds, for 8-GPU H100 systems. Those figures date from February 2025, so run the numbers again with today's rates.
Do I need a contract?
Not for on-demand: add prepaid credit, at least $5 and enough to cover the first hour, and launch. Reserved and dedicated capacity is quoted in writing for a term.
Is a GPU VPS a kind of GPUaaS?
Yes. A GPU VPS is the on-demand model in its simplest form: one virtual machine with one or more GPUs, rented by the hour.
My rule: rent on demand until you can predict the load, reserve when you can, and buy only when the hardware will stay busy and you want to run it yourself. Start with the live GPU prices, and send a capacity brief for anything that needs a term.