QuantaCloud builds dedicated GPU servers to order: each one is a single-tenant, bare-metal machine with 1 to 8 NVIDIA GPUs, configured to your spec, with full root access for your term. If you need a GPU dedicated server rather than a cloud VM, tell us the GPU, the count and the term, and we order and build the hardware.
Ask for a dedicated serverNeed a GPU today? Launch an on-demand VM. Most single-GPU VMs are running in about 3 minutes (median).
Dedicated server or on-demand VM#
The difference comes down to tenancy, access and time.
| Dedicated server | On-demand VM | |
|---|---|---|
| Tenancy | Single tenant, bare metal | A virtual machine |
| GPUs | Any NVIDIA GPU, 1 to 8 per server | Catalog GPUs, 1, 2, 4 or 8 per VM |
| Access | Full root | SSH as ubuntu |
| Operating system | The one you specify in the brief | Ubuntu 22.04 with the NVIDIA driver and Docker |
| Storage | Local NVMe and any shared storage you specify | A fixed disk that is deleted when you stop |
| Start | The lead time in your quote | Most single-GPU VMs are running in about 3 minutes (median) |
| Price | A written quote for the term | Hourly rate, first hour charged at launch, unused seconds refunded on stop |
| Region | US by default | US regions in Virginia and the Midwest |
Watch one naming trap: the on-demand template called Bare Metal is a plain Ubuntu 22.04 image with the NVIDIA driver and Docker, running on a VM. It suits most container work, and the GPU VPS page covers it. Bare metal in the hardware sense, with no hypervisor between your operating system and the GPUs, is what a dedicated server gives you.
GPUs for a dedicated server#
Any NVIDIA GPU can go in, and the choice starts with memory per GPU.
| GPU | Memory per GPU | Link between GPUs | Form factor | Page |
|---|---|---|---|---|
| RTX A6000 | 48 GB GDDR6 | NVLink bridge for pairs, 112.5 GB/s | PCIe, active cooling | RTX A6000 |
| RTX 6000 Ada | 48 GB GDDR6 | PCIe only | PCIe, active cooling | RTX 6000 Ada |
| L40 | 48 GB GDDR6 | PCIe only | PCIe, passive | L40 |
| L40S | 48 GB GDDR6 | PCIe only | PCIe, passive, 350 W | L40S |
| RTX PRO 6000 Blackwell Server Edition | 96 GB GDDR7 | PCIe only | PCIe, passive, up to 600 W | RTX PRO 6000 |
| A100 80GB | 80 GB HBM2e | SXM4: NVLink 600 GB/s. PCIe: bridge for pairs, 600 GB/s | SXM4 on HGX, or PCIe | A100 |
| H100 | 80 GB, HBM3 on SXM and HBM2e on PCIe (H100 NVL: 94 GB HBM3) | SXM: NVLink 900 GB/s, switched. PCIe: bridge for pairs, 600 GB/s | SXM on HGX, or PCIe | H100 |
| H200 | 141 GB HBM3e | SXM: NVLink 900 GB/s, switched. NVL: 2- or 4-way bridge, 900 GB/s | SXM on HGX, or PCIe (H200 NVL) | H200 |
| B200 | 180 GB HBM3E | NVLink 1.8 TB/s, switched | SXM, 8 per HGX board | B200 |
| B300 | 270 GB HBM3E | NVLink 1.8 TB/s, switched | SXM, 8 per HGX board | B300 |
PCIe or SXM is the second decision. PCIe cards fit 1 to 8 per server, talk to each other over PCIe, and get NVLink only through bridges: pairs for the A100, H100 and RTX A6000, and up to four cards for the H200 NVL. SXM GPUs come on HGX boards, and on the 8-GPU boards NVSwitch gives every GPU full NVLink bandwidth to every other. If you want Blackwell with fewer than 8 GPUs in a server, look at the RTX PRO 6000 Blackwell Server Edition: it is a PCIe card, so the count is up to you. Rack-scale systems such as the GB200 NVL72 and GB300 NVL72 are a different class, with 72 GPUs in one NVLink domain.
Most of the PCIe cards in the table are passively cooled and depend on the server's fans: NVIDIA's brief for the RTX PRO 6000 Server Edition asks for 43 to 125 CFM of airflow, depending on inlet temperature. That airflow is part of what building to spec means.
What the spec covers#
The GPUs are only part of the spec. CPU cores, system memory, local NVMe, network ports and the operating system all belong in the brief. As a reference point, NVIDIA's own 8-GPU DGX B200 pairs its GPUs with two Xeon Platinum 8570 processors (112 cores in total), 2 TB of system memory (configurable to 4 TB), 8 x 3.84 TB of NVMe for data, and a maximum power draw of about 14.3 kW. I use NVIDIA's systems as a check on balance, not as a template: if your data loader needs more CPU cores or your dataset needs more local NVMe, put the number in the brief.
Full root means you control the software on the machine: drivers, the CUDA version, containers, and GPU settings such as MIG, which splits an A100 or H100 into up to 7 isolated instances and an RTX PRO 6000 Server Edition into up to 4 instances of 24 GB.
What to send and what comes back#
A dedicated-server brief covers eight points: the GPU model and how many per server, how many servers, CPU and memory if you have a preference, local NVMe capacity, the network you need (public bandwidth, and private links if you run several servers), the operating system, the start date and term, and a monthly budget range. If one job has to span several servers, you need a fabric between them, and GPU clusters covers that case. Add your security requirements as well: QuantaCloud is not SOC 2 or HIPAA certified, so state up front what your review needs.
What comes back is the configuration, the lead time and the commercial terms, in writing, before you commit. Reserved capacity walks through the four steps, and the process is the same for one server as for a cluster.
Start on demand while we build#
Start today on the closest on-demand match, then move your work at handover.
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Prices checked 6 Oct 2026, 01:40 UTC
These are VMs: the first hour is charged at launch and unused seconds are refunded when you stop, as pricing explains, and the deploy guide covers launching one from the console. Stopping one also deletes its disk, so keep your data somewhere you control.
Dedicated GPU server FAQ#
Is a dedicated GPU server bare metal?
Yes. The server is single-tenant, and there is no hypervisor between your operating system and the hardware. QuantaCloud's on-demand instances are VMs, including the template named Bare Metal.
Do I get root access?
Yes, full root on the server for the term.
Can I get a server with one GPU?
Yes. A dedicated server can have anywhere from 1 to 8 GPUs. For a single GPU, a PCIe card such as the L40S or the H200 NVL is the fit, because SXM GPUs come on multi-GPU HGX boards.
How is a dedicated server priced?
By written quote for the configuration and the term: the spec is yours, so there is no list price. If you are weighing it against buying the hardware, the A100, H100 and H200 price guides compare purchase prices with hourly rental.
How long does a build take?
The lead time for your exact configuration comes back in writing with the quote, before you commit.
Ask for a dedicated server#
My rule: a dedicated server makes sense when you would otherwise run the same GPUs every day for months, or when you need root, your own operating system or single tenancy to pass a security review. For anything shorter, rent on demand and stop when you are done. When the dedicated case holds, send a brief with the GPU, the count, the spec and the term.