On-demand GPU

NVIDIA H100 PCIe

80 GB. Room to build.

H100 rental by the hour: the H100 PCIe with 80 GB HBM2e and FP8, 1 or 2 per VM, live prices, PCIe vs SXM specs, and what fits in 80 GB.

Full SSH accessYour choice of template
Starting from$2.59/ GPU-hr
Launch H100 PCIe
Available nowSee configurations
NVIDIA H100 PCIeHOPPER
VRAMVRAMVRAMVRAMVRAMVRAMHOPPER
Room for the work ahead.80 GB HBM2e
GPU memory
80 GB HBM2e
Memory bandwidth
2.0 TB/s
Architecture
Hopper
Deployment
On demand

Your GPU. Your configuration.

Choose where you start.

1 configuration
1 GPULowest price

1x H100 PCIe

Midwest · us-midwest-2

vCPU
20
RAM
128 GB
Disk
1,250 GB
$2.59/ instance-hour
Available nowLaunch configuration

Prices checked

The H100 you can rent on QuantaCloud by the hour is the H100 PCIe: NVIDIA's Hopper GPU with 80 GB of HBM2e and FP8 Tensor Cores, one or two per VM, from $2.59/GPU-hr. Rent it for inference and fine-tuning that fit in 80 GB. The SXM version, which comes on 4- and 8-GPU HGX boards, is not sold by the hour here: we build HGX H100 servers and clusters to order as reserved capacity.

On 2026-09-27 the catalog listed 1x and 2x H100 PCIe VMs in us-midwest-2 (Midwest). The table above is the live list, and it changes with capacity.

H100 PCIe vs H100 SXM#

The PCIe card is a smaller H100: 114 streaming multiprocessors instead of 132, HBM2e instead of HBM3, and half the power. Both are Hopper GPUs with compute capability 9.0 and 80 GB of memory, so the same code and the same FP8 kernels run on either.

SpecH100 PCIe (on demand)H100 SXM (built to order)
Streaming multiprocessors114132
GPU memory80 GB HBM2e80 GB HBM3
Memory bandwidth2.0 TB/s3.35 TB/s
FP8 Tensor Core, dense / sparse1,513 / 3,026 TFLOPS1,979 / 3,958 TFLOPS
BF16 and FP16 Tensor Core, dense / sparse756 / 1,513 TFLOPS989 / 1,979 TFLOPS
TF32 Tensor Core, dense / sparse378 / 756 TFLOPS495 / 989 TFLOPS
FP64 and FP64 Tensor Core25.6 and 51.2 TFLOPS33.5 and 66.9 TFLOPS
FP3251.2 TFLOPS66.9 TFLOPS
Max power350 WUp to 700 W
GPU to GPUNVLink bridge to one adjacent card, 600 GB/sNVLink, 900 GB/s per GPU
Host linkPCIe Gen5 x16, 128 GB/sPCIe Gen5, 128 GB/s
Form factorDual-slot, full height, full length, passive coolingSXM5 module on an HGX board
Multi-Instance GPUUp to 7 x 10 GBUp to 7 x 10 GB

On paper the SXM has 1.7 times the memory bandwidth and 1.3 times the tensor throughput (our calculation: 3.35 / 2.0 = 1.68 and 1,979 / 1,513 = 1.31). NVIDIA's own summary, in its H100 whitepaper, is that a single H100 PCIe delivers 65 percent of the SXM5's performance across ten data analytics, AI and HPC applications while drawing half the power.

The difference that decides multi-GPU work is the link. An H100 PCIe bridges to exactly one neighbour at 600 GB/s. On an 8-GPU HGX H100 board, NVSwitch gives every pair of GPUs the full 900 GB/s. My rule: jobs that fit on one or two GPUs go on the PCIe card, and training across eight goes on HGX H100 hardware, which we build to order: send a capacity brief. The on-demand card is also not the H100 NVL, which is a different PCIe model with 94 GB.

What fits in 80 GB#

The working budget is 73.3 GiB: vLLM claims 92 percent of GPU memory by default, the card reports 81,559 MiB, and weights, KV cache and activations share the budget (our calculation: 81,559 / 1,024 x 0.92).

WorkloadMemory it needsOn one H100 PCIeBasis
gpt-oss-120b (MXFP4)65.2 GB (60.8 GiB) of weights, plus 4.5 GiB of KV cache per full 131,072-token sequenceYes. OpenAI sizes it to "fit into a single 80GB GPU", and the budget leaves room for 2.8 full-length sequences: (73.3 - 60.8) / 4.5Published, our calculation
gpt-oss-20b"within 16GB of memory"Yes, with most of the card left for contextPublished
Qwen3-32B at FP834.3 GB (32.0 GiB) checkpoint, plus 8 GiB of BF16 KV cache per 32k sequenceYes, five full 32k sequences: (73.3 - 32.0) / 8 = 5.2Published size, our calculation
Qwen3-32B at BF16, one 32k sequence61.0 + 8 = 69.0 GiBJust inside the budget at default settings, with 4.3 GiB to spare before activations, so FP8 is the safer choiceOur calculation
Gemma 4 31B at BF1662.5 GB of weightsYes. Google says BF16 will "fit efficiently on a single 80GB NVIDIA H100 GPU"Published
Llama-3.3-70B at 4-bit (AWQ)39.8 GB (37.0 GiB) checkpointYes, with about 119,000 tokens of BF16 KV cache: (73.3 - 37.0) GiB / 320 KiBPublished size, our calculation
Llama-3.3-70B at FP872.7 GB (67.7 GiB) checkpointOnly at short contexts. That leaves about 5.6 GiB for KV cache, about 18,000 tokens, so for long ones use 2x H100 PCIe or one H200 NVLPublished size, our calculation
QLoRA fine-tune, 70B model41 GB (Unsloth), 40 to 48 GB (Axolotl)Yes. Unsloth reached 89,389 tokens of context for Llama 3.3 70B on 80 GBPublished
16-bit LoRA, 32B model76 GB (Unsloth), 64 to 80 GB (Axolotl)TightPublished

FP8 is where Hopper pays off at 80 GB: a 32B model at FP8 takes about half the memory of BF16, and on Hopper vLLM runs it on the FP8 Tensor Cores whether the checkpoint is per-channel or block-scaled, like Qwen's own Qwen3-32B-FP8. The L40S gets FP8 math only from per-channel checkpoints. The KV cache guide explains the per-token figures, and gpt-oss GPU requirements covers the 120b model in detail.

When to pick a different GPU#

The H100 PCIe is the wrong rental in two cases: the job needs more than about 73 GiB on one GPU, or it never uses FP8 or Hopper's extra compute.

GPUMemoryBandwidthFP8Pick it over the H100 PCIe when
H200 NVL141 GB HBM3e4.8 TB/sYesThe model and its context need more than 73 GiB, or you want 2.4 times the bandwidth
RTX PRO 6000 Blackwell96 GB GDDR71,597 or 1,792 GB/s, by editionYes, plus FP4You need 81 to 96 GB on one card, or you serve NVFP4 checkpoints
A100 80GB80 GB HBM2e1,935 GB/s (PCIe) or 2,039 GB/s (SXM4)NoThe work is BF16 and the hourly price matters more than FP8
L40S48 GB GDDR6864 GB/sYesThe job fits in 48 GB, such as a 32B model at FP8 with short contexts, or image generation
GPUMemoryFromAvailable now
H100 PCIe80 GB$2.59/GPU-hrYes
H200 NVL-Not listedNo
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
A100 PCIe 80GB80 GB$1.48/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes

The A100 comparison is closer than the names suggest. Single-user LLM decoding is bound by memory bandwidth, and on that line the H100 PCIe's 2.0 TB/s is level with the A100 80GB PCIe at 1,935 GB/s and below the A100 SXM4 at 2,039 GB/s. The H100 earns its premium on compute: 756 dense BF16 TFLOPS against the A100's 312, which is 2.4 times as much (our calculation: 756 / 312), plus FP8 and Hopper-only kernels such as FlashAttention-3, which lists the H100 as a requirement. So I put prefill-heavy serving, FP8 models and training on the H100 PCIe, and BF16 decoding on a budget on the A100. L40S vs H100 and RTX PRO 6000 vs H100 go through two more of the table's rows in detail.

If you are weighing a purchase instead, the H100 price guide works through buying against renting, and if the alternative is a desktop card, H100 vs RTX 5090 covers when 80 GB matters. For eight SXM GPUs or more, the GPU clusters page covers multi-node builds, and InfiniBand vs NVLink explains which link carries what.

Measured on QuantaCloud#

Across QuantaCloud, most single-GPU VMs are running in about 3 minutes (median). The GPU0 to GPU1 entry in nvidia-smi topo -m decides how you split a model over the 2x VM: an NV entry means the bridge is in place and tensor parallelism is the usual choice, while any other entry, such as PHB, NODE or SYS, means PCIe only, where vLLM recommends pipeline parallelism. Check the NVIDIA driver version too, since vLLM's default image is a CUDA 13 build that needs an R580 or newer driver.

Launch it with a template#

The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.

TemplateWhat I would run on an H100 PCIeLaunch
Bare Metal: Ubuntu 22.04, NVIDIA driver, DockervLLM in Docker for gpt-oss-120b or FP8 models, following the vLLM Docker guideLaunch Bare Metal
PyTorch + JupyterQLoRA on a 70B model, or FP8 training experiments, from a Jupyter notebookLaunch PyTorch + Jupyter
Open WebUI + OllamaA private chat on gpt-oss-120b or Gemma 4 31BLaunch Open WebUI
ComfyUIFLUX.2 [dev], whose fp8 transformer (35.5 GB) and fp8 text encoder (18 GB) fit on the card together. Check its non-commercial licence before you serve itLaunch ComfyUI

Every instance is a VM: Bare Metal is the name of the plain Ubuntu template, not bare-metal hardware. SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login. Download the model weights you need after boot. The templates docs list what each one contains.

FAQ#

Is this the H100 SXM or the H100 PCIe?

The PCIe version: 80 GB of HBM2e at 2.0 TB/s with a 350 W limit. The SXM version is available as a reserved build.

Does gpt-oss-120b run on one H100 PCIe?

Yes. OpenAI sizes gpt-oss-120b for a single 80 GB GPU, and its MXFP4 weights take about 65 GB. That leaves room for two full 131k-token conversations and most of a third at vLLM's default memory setting, so for many long conversations at once I would move to the H200 NVL.

Only the topology check can say. The H100 PCIe supports a bridge to one neighbour, but QuantaCloud's catalog lists these offers as PCIe. Unless the GPU0 to GPU1 entry in nvidia-smi topo -m starts with NV, the two GPUs talk over PCIe.

Can I rent eight H100s or an H100 cluster?

Not by the hour. Eight H100s in one server is the HGX H100 platform, and QuantaCloud builds HGX H100 servers and InfiniBand clusters to order as reserved capacity, with the configuration, lead time and terms confirmed in writing. B200 vs H100 covers what Blackwell changes.

Should I buy an H100 instead of renting one?

The answer turns on how many hours a year you will keep it busy. The H100 price guide works through purchase prices against hourly rates.

What happens when I stop the instance?

Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots, so copy results off first. The unused seconds of the current hour are refunded to your balance. Pricing and the billing docs have the full rules.

Is it available right now?

The table at the top is live. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 the H100 PCIe was in us-midwest-2.


The rule I follow: rent the H100 PCIe when the job fits in 80 GB and uses what Hopper adds, meaning FP8 or the extra compute. If it needs more memory, go to the H200 NVL. If it is BF16 decoding that fits in 80 GB, check the A100's live price in the GPU catalog first, since its memory bandwidth is about the same. Start with a 1x VM and the vLLM Docker guide, and measure your own tokens per second before you add a second card. Launch H100 PCIe

Keep building

Choose your next step.