Reserved capacity is built to order: you tell us the GPUs, the count, the network and the term, and we order and build the hardware as dedicated capacity for your team. If you are looking for a GPU lease, this is how it works at QuantaCloud: hardware made to your spec and held for an agreed term, with the configuration, lead time and commercial terms in writing before you commit.
Send a capacity briefNeed a GPU today? Launch one on demand. QuantaCloud GPU instances run in US regions with 1, 2, 4 or 8 GPUs per VM, and most single-GPU VMs are running in about 3 minutes (median).
What we can build#
Any NVIDIA GPU configuration can be built, from one server to an InfiniBand cluster: we order the hardware and build it to your spec. Each build takes one of three shapes.
| Shape | What you get | More detail |
|---|---|---|
| Dedicated servers | Single-tenant, bare-metal servers with 1 to 8 GPUs each and full root access | Dedicated GPU servers |
| Multi-node clusters | GPU servers joined by an InfiniBand or RoCE fabric, with NVLink inside each server where the platform has it | GPU clusters |
| Rack-scale NVLink | GB200 NVL72 or GB300 NVL72: 72 GPUs in one NVLink domain at 130 TB/s, liquid-cooled. We discuss power and cooling with you before quoting | GB200 and GB300 |
The GPU list starts with the on-demand catalog and goes past it. You can reserve anything the catalog carries, such as the RTX A6000 and L40S at 48 GB, the A100 at 80 GB or the RTX PRO 6000 Blackwell at 96 GB, and GPUs it does not carry, such as the H100 and H200 in SXM form on HGX boards. The B200 and B300 are reserved builds only: 8 GPUs per HGX board, 180 GB and 270 GB per GPU, and NVLink at 1.8 TB/s per GPU. B300 vs B200 covers the choice between those two. H200 vs B200 and B200 vs H100 compare the B200 with the H200 and the H100. Storage, region and term are part of the spec as well, and builds are in the US by default.
How it works#
Every build follows the same four steps.
- You send a capacity brief with the form at the end of this page.
- We reply with a proposed configuration, a lead time and commercial terms, all in writing.
- You review the quote. Once you accept it, we order the hardware and build it to the agreed configuration.
- At handover you get access to your servers for the term.
What to put in your brief#
A useful brief answers ten questions. If you do not know an answer yet, say so: a stated gap is more useful than a guess, because the quote is built on what you write.
| Question | What to tell us |
|---|---|
| GPU and count | The GPU model and total count. If you are unsure, the model size and precision you need to fit |
| Workload | Training, fine-tuning or inference, and the framework, such as PyTorch FSDP, DeepSpeed or vLLM |
| Interconnect | Whether one job spans several servers. If it does, InfiniBand or RoCE, and the port speed per GPU (400 or 800 Gb/s) |
| Storage | Capacity in TB, local NVMe or shared, and how fast you need to read data and write checkpoints |
| Software | The operating system and scheduler you plan to run, such as Slurm or Kubernetes |
| Region | US by default. Say so if your data has to stay in a particular place |
| Start date | When you need the capacity, and what a later date would cost you |
| Term | How long you need it |
| Budget | A monthly range, even a rough one |
| Security and compliance | Anything your auditors or your customers will ask about |
Reserved or on-demand#
The rule I follow is simple: reserve when your plan depends on the hardware being there on a date, and stay on demand while it does not.
| On-demand | Reserved | |
|---|---|---|
| Availability | Whatever the live catalog shows when you launch | Built for you and held for the term |
| Commitment | None: stop when you are done | The term in your quote |
| Price basis | Hourly rate. The first hour is charged at launch and unused seconds are refunded when you stop | A written quote for the configuration and term |
| Setup time | Most single-GPU VMs are running in about 3 minutes (median) | The lead time in your quote |
| Hardware | VMs with 1, 2, 4 or 8 GPUs in US regions | Any NVIDIA GPU, from one bare-metal server to a cluster |
| Storage | A fixed disk that is deleted when you stop | Specified in your brief |
Reserved and on-demand work together: while your build is under way, you can develop and test on an on-demand VM from the GPU catalog, and the deploy guide walks through the console. Copy your results off before you stop it: stopping an on-demand instance terminates it and deletes its disk, and there are no volumes or snapshots to fall back on. Pricing covers the on-demand billing rules, and the H100 price guide compares buying an H100 with renting one by the hour.
Lead times#
Lead times depend on GPU availability at the time of your order. The one for your exact configuration comes back in writing with the quote, before you commit, and a change to the GPU model or the count can move it.
If you need GPUs this week, start on demand and send the brief at the same time.
Security and data#
QuantaCloud is not SOC 2 or HIPAA certified. Put your security and compliance requirements in the brief up front: where data may live, who may access the servers and which audits apply.
Reserved capacity FAQ#
Can we start on demand while you build?
Yes. Launch on-demand VMs from the catalog today and move your work to the reserved build at handover. Keep your own copies of code, data and checkpoints, because stopping an on-demand instance deletes its disk.
Can you build InfiniBand clusters?
Yes, as part of a build. Multi-node builds use InfiniBand or RoCE between servers, and NVLink inside each server where the platform has it. GPU clusters covers node platforms and fabric choices, and InfiniBand vs NVLink explains which link carries which traffic.
How is a reserved build billed?
By the terms in your quote. The price, the payment terms and the length of the term are set out in writing for your configuration. On-demand billing works differently, and pricing explains it.
Is there an SLA?
QuantaCloud does not publish an SLA or an uptime commitment. If your procurement needs service terms, list them in the brief.
Send a capacity brief#
My rule of thumb: if a late GPU would cost you more than a term commitment, send the brief now. If it would not, start on demand today and send the brief once the plan is firm. Either way, what comes back is in writing, and you decide from there.