QuantaCloud rents the NVIDIA A100 80GB in two versions: the SXM4, one or eight per VM, and the PCIe card, two or four per VM, from $1.48/GPU-hr. An A100 rental gets you the same 80 GB per GPU as the H100 PCIe without FP8, which is the trade to weigh first. The 8x SXM4 VM puts 640 GB of GPU memory in one machine, the most of any configuration in the catalog on 2026-09-27, for fine-tuning jobs that need eight GPUs.
On 2026-09-27 the catalog listed the SXM4 as 1x in us-east-1 (Virginia), us-midwest-1 and us-midwest-2, and as 8x in us-midwest-2. The PCIe card was listed as 2x and 4x in us-midwest-2, and a 1x A100 PCIe appeared 13 minutes later. The table above is the live list.
A100 SXM4 vs A100 PCIe#
The chip is the same, and NVIDIA publishes the same peak throughput for both versions. What differs is the power limit, a little memory bandwidth, and how the GPUs connect.
| Spec | A100 SXM4 80GB | A100 PCIe 80GB |
|---|---|---|
| Architecture | Ampere, compute capability 8.0 | Ampere, compute capability 8.0 |
| GPU memory | 80 GB HBM2e | 80 GB HBM2e |
| Memory bandwidth | 2,039 GB/s | 1,935 GB/s |
| BF16 and FP16 Tensor Core, dense / sparse | 312 / 624 TFLOPS | 312 / 624 TFLOPS |
| TF32 Tensor Core, dense / sparse | 156 / 312 TFLOPS | 156 / 312 TFLOPS |
| INT8 Tensor Core, dense / sparse | 624 / 1,248 TOPS | 624 / 1,248 TOPS |
| FP64 and FP64 Tensor Core | 9.7 and 19.5 TFLOPS | 9.7 and 19.5 TFLOPS |
| FP32 | 19.5 TFLOPS | 19.5 TFLOPS |
| FP8 and FP4 | No | No |
| Max power | 400 W | 300 W |
| GPU to GPU | NVLink at 600 GB/s through the HGX A100 server board | NVLink bridge for 2 GPUs, 600 GB/s |
| Host link | PCIe Gen4, 64 GB/s | PCIe Gen4, 64 GB/s |
| Multi-Instance GPU | Up to 7 x 10 GB | Up to 7 x 10 GB |
The NVLink row is the one that changes your choice. QuantaCloud's API marks the SXM4 offers as SXM with NVLink and the PCIe offers as PCIe. On a 1x VM the flag changes nothing, because there is no second GPU to talk to. On the 8x VM it is the reason to choose the SXM4: multi-GPU training exchanges gradients between GPUs on every step, and NVLink carries that at 600 GB/s where PCIe Gen4 manages 64 GB/s. I do not take the flag on trust: nvidia-smi topo -m on the VM has to show NV entries between GPU pairs.
What fits in 80 GB, and in 640 GB#
The working budget on one A100 is 73.6 GiB: vLLM claims 92 percent of GPU memory by default, the card reports 81,920 MiB, and weights, KV cache and activations share the budget (our calculation: 81,920 / 1,024 x 0.92).
| Workload | Memory it needs | On A100 80GB | Basis |
|---|---|---|---|
| gpt-oss-120b (MXFP4) | 65.2 GB (60.8 GiB) of weights, plus 4.5 GiB of KV cache per full 131,072-token sequence | Yes on 1x. vLLM runs MXFP4 on the A100 through its Marlin kernels, with room for 2.8 full-length sequences: (73.6 - 60.8) / 4.5 | Published, our calculation |
| Llama-3.3-70B at 4-bit (AWQ) | 39.8 GB (37.0 GiB) checkpoint | Yes on 1x, with about 120,000 tokens of BF16 KV cache: (73.6 - 37.0) GiB / 320 KiB | Published size, our calculation |
| Llama-3.3-70B at FP8 | 72.7 GB (67.7 GiB) checkpoint, run weight-only on Ampere | Only at short contexts on 1x, where about 5.9 GiB, about 19,000 tokens, would be left for KV cache. Yes on 2x PCIe | Published size, our calculation |
| Llama-3.3-70B at BF16 | 141.1 GB (131.4 GiB) of weights | Yes on 4x PCIe (294.4 GiB budget) or 8x SXM4 | Our calculation |
| Qwen3-32B at BF16, one 32k sequence | 61.0 + 8 = 69.0 GiB | Just inside the budget on 1x, with 4.6 GiB to spare before activations. Comfortable on 2x PCIe | Our calculation |
| QLoRA fine-tune, 70B model | 41 GB (Unsloth), 40 to 48 GB (Axolotl) | Yes on 1x | Published |
| 16-bit LoRA, 32B model | 76 GB (Unsloth), 64 to 80 GB (Axolotl) | Tight on 1x. Fits across 2x PCIe with FSDP | Published |
| 16-bit LoRA, 70B model | 164 GB (Unsloth), 2x 80 GB (Axolotl) | 4x PCIe or 8x SXM4 | Published |
| Full fine-tune, Qwen3-8B (8.19B parameters) | 131 GB of model states at 16 bytes per parameter: 8.19 x 16 | 8x SXM4 with ZeRO-3 or FSDP: 16.4 GB of model states per GPU (131 / 8), plus activations | Our calculation |
| Full fine-tune, 70B model | 1,120 to 1,260 GB of model states (70 x 16 to 70 x 18) | No. It needs at least 16x 80 GB, which is a multi-node build | Our calculation |
The 16 bytes per parameter is mixed-precision Adam as the ZeRO paper counts it: 2 for 16-bit weights, 2 for 16-bit gradients and 12 for the FP32 master weights and Adam states. Hugging Face counts 18 when gradients stay in FP32. The LoRA, QLoRA and full fine-tuning VRAM guide covers the rest, and Unsloth fine-tuning is a worked QLoRA run.
Where the A100 falls short#
The A100's real gap is FP8. FP8 Tensor Core math needs compute capability 8.9 or higher, and the A100 is 8.0. vLLM still loads FP8 checkpoints on an A100, but runs them weight-only (W8A16) through its Marlin kernels, so you get the memory saving without the FP8 speed. ComfyUI makes the same check and only computes in FP8 on 8.9 and up. The second gap is compute: 312 dense BF16 TFLOPS against 756 on the H100 PCIe, which is 2.4 times as much (our calculation: 756 / 312).
| GPU | Memory | Bandwidth | FP8 | Pick it over the A100 when |
|---|---|---|---|---|
| H100 PCIe | 80 GB HBM2e | 2.0 TB/s | Yes | You want FP8 or 2.4 times the BF16 compute at the same 80 GB |
| H200 NVL | 141 GB HBM3e | 4.8 TB/s | Yes | The model needs more than 80 GB on one GPU |
| L40S | 48 GB GDDR6 | 864 GB/s | Yes | The job fits in 48 GB and can use FP8 math, such as image generation or a 32B model in a per-channel FP8 checkpoint |
| RTX A6000 | 48 GB GDDR6 | 768 GB/s | No | The job fits in 48 GB and the hourly price matters most |
| GPU | Memory | From | Available now |
|---|---|---|---|
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
Where the A100 still wins is memory per dollar. On 2026-09-27 a 1x A100 SXM4 cost $1.50 an hour against $2.59 for a 1x H100 PCIe, 42 percent less for the same 80 GB and slightly more bandwidth (our calculation: 1 - 1.50 / 2.59 = 0.42). For BF16 decoding, which is bound by memory bandwidth, the two are closer than their names suggest. The L40S vs A100 and RTX A6000 vs A100 comparisons cover the 48 GB alternatives, A100 vs RTX 5090 covers the desktop card, and the A100 price guide works through buying against renting. A100 vs H200 NVL asks whether the upgrade is worth it, and RTX PRO 6000 vs A100 compares it with the RTX PRO 6000 Blackwell.
Past eight GPUs, the step is a reserved build. A full fine-tune of a 70B model needs a multi-node cluster, and QuantaCloud builds GPU clusters to order: send a capacity brief.
Measured on QuantaCloud#
Across QuantaCloud, most single-GPU VMs are running in about 3 minutes (median), and 8-GPU VMs take about 10. Check the NVIDIA driver version before you pull an image: vLLM's default image is a CUDA 13 build that needs an R580 or newer driver.
Launch it with a template#
The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.
| Template | What I would run on an A100 | Launch |
|---|---|---|
| Bare Metal: Ubuntu 22.04, NVIDIA driver, Docker | vLLM in Docker for gpt-oss-120b or a 4-bit 70B model. On the 8x VM, a PyTorch container running torchrun | Launch Bare Metal |
| PyTorch + Jupyter | QLoRA on a 70B model, or 16-bit LoRA up to 32B, from a Jupyter notebook | Launch PyTorch + Jupyter |
| Open WebUI + Ollama | A private chat on gpt-oss-120b | Launch Open WebUI |
| ComfyUI | Wan 2.2 A14B video at 720p, which peaked at 59.8 GB in Wan's own single-GPU test. On Ampere, fp8 files save memory but compute runs in 16-bit | Launch ComfyUI |
Every instance is a VM: Bare Metal is the name of the plain Ubuntu template, not bare-metal hardware. SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login. Download the model weights you need after boot. The templates docs list what each one contains.
FAQ#
Should I pick the A100 SXM4 or the A100 PCIe?
For one GPU, take whichever is cheaper when you check: the chip, the 80 GB and the published TFLOPS are the same. For eight GPUs training together, take the 8x SXM4: SXM4 GPUs are built to exchange gradients over NVLink, and nvidia-smi topo -m shows whether the VM exposes it. The PCIe 2x and 4x VMs suit inference and independent jobs, where each GPU mostly works on its own.
Is NVLink working on the 8x A100 SXM4 VM?
The API flags it. Run nvidia-smi topo -m on your own VM to check. An entry of NV plus a number between two GPUs means NVLink, and the number counts the links.
Does the A100 support FP8?
No. FP8 Tensor Cores start at compute capability 8.9, and the A100 is 8.0. FP8 checkpoints still load in vLLM and halve the weight memory, but the math runs in 16-bit. If FP8 speed matters, rent the H100 PCIe, or the L40S with a per-channel FP8 checkpoint.
Can I run gpt-oss-120b on an A100?
Yes. vLLM's MXFP4 path needs compute capability 8.0, which the A100 meets through the Marlin kernels, and the 65 GB of weights fit in 80 GB. For many long conversations at once, the extra memory of the H200 NVL is the better fit.
Is it available right now?
The table at the top is live. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 the A100 was in us-east-1 (Virginia), us-midwest-1 and us-midwest-2.
What happens when I stop the instance?
Stopping terminates the VM and deletes its disk, and there are no volumes or snapshots. On a long fine-tune, copy checkpoints off the VM as you go. The unused seconds of the current hour are refunded to your balance. Pricing and the billing docs have the full rules.
The rule I follow for the A100: rent it when the job needs 80 GB per GPU and BF16 is good enough, and rent the 8x SXM4 VM when a fine-tune needs eight GPUs in one machine. If the model needs FP8 or more than 80 GB on one GPU, step up to the H100 PCIe or the H200 NVL. If it fits in 48 GB, check the L40S and the RTX A6000 in the GPU catalog first. Start on a 1x SXM4 to confirm the job fits, then run the same code on the 8x VM. Launch A100 SXM4