Dedicated capacity

Dedicated GPUs, built to your spec.

Dedicated GPU servers built to your spec: single-tenant bare metal with 1 to 8 NVIDIA GPUs and full root. Send a brief, get the quote in writing.

QuantaCloud builds dedicated GPU servers to order: each one is a single-tenant, bare-metal machine with 1 to 8 NVIDIA GPUs, configured to your spec, with full root access for your term. If you need a GPU dedicated server rather than a cloud VM, tell us the GPU, the count and the term, and we order and build the hardware.

Ask for a dedicated server

Need a GPU today? Launch an on-demand VM. Most single-GPU VMs are running in about 3 minutes (median).

Dedicated server or on-demand VM#

The difference comes down to tenancy, access and time.

Dedicated serverOn-demand VM
TenancySingle tenant, bare metalA virtual machine
GPUsAny NVIDIA GPU, 1 to 8 per serverCatalog GPUs, 1, 2, 4 or 8 per VM
AccessFull rootSSH as ubuntu
Operating systemThe one you specify in the briefUbuntu 22.04 with the NVIDIA driver and Docker
StorageLocal NVMe and any shared storage you specifyA fixed disk that is deleted when you stop
StartThe lead time in your quoteMost single-GPU VMs are running in about 3 minutes (median)
PriceA written quote for the termHourly rate, first hour charged at launch, unused seconds refunded on stop
RegionUS by defaultUS regions in Virginia and the Midwest

Watch one naming trap: the on-demand template called Bare Metal is a plain Ubuntu 22.04 image with the NVIDIA driver and Docker, running on a VM. It suits most container work, and the GPU VPS page covers it. Bare metal in the hardware sense, with no hypervisor between your operating system and the GPUs, is what a dedicated server gives you.

GPUs for a dedicated server#

Any NVIDIA GPU can go in, and the choice starts with memory per GPU.

GPUMemory per GPULink between GPUsForm factorPage
RTX A600048 GB GDDR6NVLink bridge for pairs, 112.5 GB/sPCIe, active coolingRTX A6000
RTX 6000 Ada48 GB GDDR6PCIe onlyPCIe, active coolingRTX 6000 Ada
L4048 GB GDDR6PCIe onlyPCIe, passiveL40
L40S48 GB GDDR6PCIe onlyPCIe, passive, 350 WL40S
RTX PRO 6000 Blackwell Server Edition96 GB GDDR7PCIe onlyPCIe, passive, up to 600 WRTX PRO 6000
A100 80GB80 GB HBM2eSXM4: NVLink 600 GB/s. PCIe: bridge for pairs, 600 GB/sSXM4 on HGX, or PCIeA100
H10080 GB, HBM3 on SXM and HBM2e on PCIe (H100 NVL: 94 GB HBM3)SXM: NVLink 900 GB/s, switched. PCIe: bridge for pairs, 600 GB/sSXM on HGX, or PCIeH100
H200141 GB HBM3eSXM: NVLink 900 GB/s, switched. NVL: 2- or 4-way bridge, 900 GB/sSXM on HGX, or PCIe (H200 NVL)H200
B200180 GB HBM3ENVLink 1.8 TB/s, switchedSXM, 8 per HGX boardB200
B300270 GB HBM3ENVLink 1.8 TB/s, switchedSXM, 8 per HGX boardB300

PCIe or SXM is the second decision. PCIe cards fit 1 to 8 per server, talk to each other over PCIe, and get NVLink only through bridges: pairs for the A100, H100 and RTX A6000, and up to four cards for the H200 NVL. SXM GPUs come on HGX boards, and on the 8-GPU boards NVSwitch gives every GPU full NVLink bandwidth to every other. If you want Blackwell with fewer than 8 GPUs in a server, look at the RTX PRO 6000 Blackwell Server Edition: it is a PCIe card, so the count is up to you. Rack-scale systems such as the GB200 NVL72 and GB300 NVL72 are a different class, with 72 GPUs in one NVLink domain.

Most of the PCIe cards in the table are passively cooled and depend on the server's fans: NVIDIA's brief for the RTX PRO 6000 Server Edition asks for 43 to 125 CFM of airflow, depending on inlet temperature. That airflow is part of what building to spec means.

What the spec covers#

The GPUs are only part of the spec. CPU cores, system memory, local NVMe, network ports and the operating system all belong in the brief. As a reference point, NVIDIA's own 8-GPU DGX B200 pairs its GPUs with two Xeon Platinum 8570 processors (112 cores in total), 2 TB of system memory (configurable to 4 TB), 8 x 3.84 TB of NVMe for data, and a maximum power draw of about 14.3 kW. I use NVIDIA's systems as a check on balance, not as a template: if your data loader needs more CPU cores or your dataset needs more local NVMe, put the number in the brief.

Full root means you control the software on the machine: drivers, the CUDA version, containers, and GPU settings such as MIG, which splits an A100 or H100 into up to 7 isolated instances and an RTX PRO 6000 Server Edition into up to 4 instances of 24 GB.

What to send and what comes back#

A dedicated-server brief covers eight points: the GPU model and how many per server, how many servers, CPU and memory if you have a preference, local NVMe capacity, the network you need (public bandwidth, and private links if you run several servers), the operating system, the start date and term, and a monthly budget range. If one job has to span several servers, you need a fabric between them, and GPU clusters covers that case. Add your security requirements as well: QuantaCloud is not SOC 2 or HIPAA certified, so state up front what your review needs.

What comes back is the configuration, the lead time and the commercial terms, in writing, before you commit. Reserved capacity walks through the four steps, and the process is the same for one server as for a cluster.

Start on demand while we build#

Start today on the closest on-demand match, then move your work at handover.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 6 Oct 2026, 01:40 UTC

These are VMs: the first hour is charged at launch and unused seconds are refunded when you stop, as pricing explains, and the deploy guide covers launching one from the console. Stopping one also deletes its disk, so keep your data somewhere you control.

Dedicated GPU server FAQ#

Is a dedicated GPU server bare metal?

Yes. The server is single-tenant, and there is no hypervisor between your operating system and the hardware. QuantaCloud's on-demand instances are VMs, including the template named Bare Metal.

Do I get root access?

Yes, full root on the server for the term.

Can I get a server with one GPU?

Yes. A dedicated server can have anywhere from 1 to 8 GPUs. For a single GPU, a PCIe card such as the L40S or the H200 NVL is the fit, because SXM GPUs come on multi-GPU HGX boards.

How is a dedicated server priced?

By written quote for the configuration and the term: the spec is yours, so there is no list price. If you are weighing it against buying the hardware, the A100, H100 and H200 price guides compare purchase prices with hourly rental.

How long does a build take?

The lead time for your exact configuration comes back in writing with the quote, before you commit.

Ask for a dedicated server#

My rule: a dedicated server makes sense when you would otherwise run the same GPUs every day for months, or when you need root, your own operating system or single tenancy to pass a security review. For anything shorter, rent on demand and stop when you are done. When the dedicated case holds, send a brief with the GPU, the count, the spec and the term.

Keep building

Choose your next step.