BF16 is the right default for training and serving on every NVIDIA GPU from the Ampere generation on: it keeps FP32's range in 16 bits, so it rarely overflows. FP16 has three more bits of precision but far less range, from about 6 x 10^-8 up to 65,504, which is why training in it needs loss scaling. FP8 halves the memory of either and doubles the Tensor Core throughput, but only GPUs with compute capability 8.9 or newer compute in it: Ada, Hopper and Blackwell. FP4 halves the memory again, and only Blackwell computes in it.
On QuantaCloud that splits the catalog in three. The RTX A6000 and A100 do BF16 and FP16 but not FP8. The L40, L40S, RTX 6000 Ada, H100 PCIe and H200 NVL add FP8. The RTX PRO 6000 Blackwell adds FP4. The GPU catalog has live prices for every one of them.
The formats side by side#
Every floating-point format splits its bits between an exponent, which sets the range, and a mantissa, which sets the precision. The step after 1.0 is the gap to the next number a format can store, and it doubles with every mantissa bit you take away (our calculation from the bit layouts):
| Format | Bits: sign, exponent, mantissa | Largest value | Step after 1.0 | Main use |
|---|---|---|---|---|
| FP32 | 1, 8, 23 | About 3.4 x 10^38 | 0.00000012 | Master weights, optimizer states, reductions |
| TF32 | 1, 8, 10 | About 3.4 x 10^38 | 0.00098 | Tensor Core math on FP32 inputs |
| BF16 | 1, 8, 7 | About 3.4 x 10^38 | 0.0078 | Training and serving by default |
| FP16 | 1, 5, 10 | 65,504 | 0.00098 | Older GPUs, and models trained in FP16 |
| FP8 E4M3 | 1, 4, 3 | 448 | 0.125 | Weights and activations |
| FP8 E5M2 | 1, 5, 2 | 57,344 | 0.25 | Gradients in FP8 training |
| FP4 E2M1 | 1, 2, 1 | 6 | 0.5 | Weights and activations, with block scales (NVFP4, MXFP4) |
Two rows explain most of the practical differences. BF16 and FP16 are the same size and make opposite trades: BF16 keeps FP32's 8-bit exponent and gives up precision, while FP16 keeps more precision and loses range. In FP16 the smallest normal number is about 6.1 x 10^-5, and NVIDIA's mixed-precision guide puts the smallest representable one at about 6.0 x 10^-8, so small gradients round to zero unless the loss is scaled up first.
The FP8 and FP4 rows only work with scale factors. With 448 as the ceiling for E4M3, and 6 for FP4, real weights and activations have to be divided by a scale before they fit, and a checkpoint stores that scale per tensor, per channel or per block of values. The NVFP4 guide shows how the 4-bit formats do it.
TF32 is not a storage format. It is what Ampere and newer Tensor Cores do with FP32 inputs: round them to 10 mantissa bits, multiply, and accumulate in FP32. PyTorch's own example on an A100 ran a large matrix multiply about 7 times faster in TF32, with a relative error of 0.0022 against 0.000039 in full FP32.
Which GPUs compute each format#
Hardware support follows the architecture, and FP8 is the dividing line in the on-demand catalog:
| GPU | Architecture, compute capability | FP16, BF16, TF32 | FP8 | FP4 |
|---|---|---|---|---|
| RTX A6000 | Ampere, 8.6 | Yes | No | No |
| A100 80GB | Ampere, 8.0 | Yes | No | No |
| L40, L40S, RTX 6000 Ada | Ada, 8.9 | Yes | Yes | No |
| H100 PCIe, H200 NVL | Hopper, 9.0 | Yes | Yes | No |
| RTX PRO 6000 Blackwell | Blackwell, 12.0 | Yes | Yes | Yes |
| B200, B300 (built to order) | Blackwell, 10.0 and 10.3 | Yes | Yes | Yes |
NVIDIA's PTX instruction set draws the same lines: FP8 matrix instructions need sm_89 or newer, and FP4 ones first appear with Blackwell, in sm_100 and sm_120. A GPU without FP8 can still load an FP8 checkpoint. vLLM runs it weight-only on Ampere, storing the weights in FP8 and doing the math in 16-bit, so you keep the memory saving and lose the speed.
Ada has one more catch in vLLM 0.30.0. Per-channel FP8 checkpoints, the kind llm-compressor's FP8_DYNAMIC scheme produces, get FP8 math on Ada and newer. Block-scaled checkpoints, such as Qwen's own Qwen3-32B-FP8, get FP8 math only on Hopper and Blackwell and run weight-only on the L40, L40S and RTX 6000 Ada.
Tensor Core throughput by format#
On every GPU in this table, each step down in bits doubles the dense Tensor Core throughput. These are NVIDIA's dense figures, without sparsity:
| GPU | BF16 and FP16 | FP8 | FP4 |
|---|---|---|---|
| A100 80GB | 312 TFLOPS | Not supported | Not supported |
| L40 | 181 TFLOPS | 362 TFLOPS | Not supported |
| L40S | 362 TFLOPS | 733 TFLOPS | Not supported |
| H100 PCIe | 756 TFLOPS | 1,513 TFLOPS | Not supported |
| H200 NVL | About 836 TFLOPS | About 1,671 TFLOPS | Not supported |
| RTX PRO 6000 Blackwell (Workstation Edition) | 503.8 TFLOPS | 1,007.6 TFLOPS | 2,015.2 TFLOPS |
| B200 (built to order) | 2,250 TFLOPS | 4,500 TFLOPS | 9,000 TFLOPS |
NVIDIA publishes the H200 NVL's Tensor Core figures with sparsity only, so its row is half of those (our calculation). The RTX PRO 6000 row comes from NVIDIA's RTX PRO Blackwell architecture whitepaper, which lists every Workstation Edition rate twice, dense and with sparsity. The round figures on NVIDIA's Server Edition page, 1 PFLOP for BF16, 2 for FP8 and 4 for FP4, match the whitepaper's sparse rates, so they are twice the dense figures above. QuantaCloud's catalog lists the card only as RTX PRO 6000 Blackwell, and nvidia-smi -q on the VM shows the edition.
TF32 runs at half the BF16 rate on the A100, L40S, H100 PCIe, H200 NVL and RTX PRO 6000: 156 dense TFLOPS on the A100, 378 on the H100 PCIe and 251.9 on the Workstation Edition, for example. The GPU comparisons set these cards side by side with live prices.
Memory: what each format costs#
Memory scales with the bits: 4 bytes per parameter in FP32, 2 in BF16 or FP16, 1 in FP8 and about half a byte in a 4-bit format. Real checkpoints come out a little larger, because the embeddings and output layer usually stay in BF16 and the scales take space. These are the file sizes on the Hugging Face API on 2026-09-28:
| Model | BF16 | FP8 | NVFP4 |
|---|---|---|---|
| Llama 3.3 70B Instruct | 141.1 GB | 72.7 GB | 42.7 GB |
| Qwen3-32B | 65.5 GB | 34.3 GB | 20.7 GB |
FP8 cuts both models to about half, 1.9 times smaller (our calculation). The KV cache has its own precision setting: vLLM's --kv-cache-dtype fp8 halves it, and without calibration it uses scales of 1.0, so check the output on your own prompts. The KV cache guide shows how much memory that frees.
Training is different. Standard mixed-precision Adam needs 16 to 18 bytes per parameter whichever format the matrix multiplies use, because the master weights and optimizer states stay in FP32. NVIDIA's Transformer Engine documentation is blunt about it: FP8 training does not always reduce memory compared to BF16. The fine-tuning VRAM guide has the numbers by model size.
When lower precision runs faster#
Lower precision runs faster where the math or the memory reads are the bottleneck, and LLM serving has one of each. NVIDIA's inference guide says reading the prompt (prefill) "effectively saturates GPU utilization", while generating tokens one at a time is "a memory-bound operation". FP8 Tensor Cores speed up the first. Smaller weights speed up the second on any GPU, because there are fewer bytes to read for every token. vLLM's documentation puts FP8 at half the model memory and up to 1.6 times the throughput.
Precision for training#
The rule I follow for training is BF16 mixed precision on anything from Ampere up, with the master weights in FP32. In PyTorch that is autocast around the forward pass, with no loss scaling:
import torch
for x, y in loader:
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
loss = loss_fn(model(x), y)
loss.backward()
optimizer.step()
optimizer.zero_grad()
FP16 needs a gradient scaler. PyTorch's documentation explains why: FP16 gradients with small magnitudes flush to zero, so the scaler multiplies the loss before the backward pass and divides the gradients again before the optimizer step:
scaler = torch.amp.GradScaler("cuda")
for x, y in loader:
with torch.autocast(device_type="cuda", dtype=torch.float16):
loss = loss_fn(model(x), y)
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
optimizer.zero_grad()
The same documentation warns that most models pretrained in BF16 cannot run in FP16's range at all: their gradients overflow past 65,504 instead of underflowing. Every GPU QuantaCloud rents supports BF16, so I would train in FP16 only a model that was trained in FP16.
If your code still runs in FP32, TF32 is a one-line speed-up for matrix multiplies, and it is off by default in PyTorch since version 1.12. PyTorch 2.9 and newer set it like this:
torch.backends.cuda.matmul.fp32_precision = "tf32"
FP8 training is for large jobs on Ada, Hopper or Blackwell. NVIDIA's Transformer Engine runs the matrix multiplies in FP8, using E4M3 in the forward pass and E5M2 for gradients by default, and keeps the weights in higher precision:
import torch
import transformer_engine.pytorch as te
from transformer_engine.common.recipe import Float8CurrentScaling, Format
recipe = Float8CurrentScaling(fp8_format=Format.HYBRID)
layer = te.Linear(1024, 1024, params_dtype=torch.bfloat16)
inp = torch.randn(32, 128, 1024, dtype=torch.bfloat16, device="cuda")
with te.autocast(enabled=True, recipe=recipe):
out = layer(inp)
out.sum().backward()
DeepSeek-V3 is the public example of FP8 training at scale. Its report puts the matrix multiplies in FP8, which in theory doubles their speed over BF16, and keeps the embeddings, output head, normalization and attention in BF16 or FP32, along with the master weights, gradients and optimizer states. On its smaller validation runs, the loss stayed within 0.25% of a BF16 baseline. For LoRA and QLoRA fine-tuning I start in BF16, and the Unsloth tutorial walks through a full QLoRA run.
Precision for inference#
Serving is where the format choice pays most, and it depends on the GPU generation:
| Generation | QuantaCloud GPUs | What I would serve in |
|---|---|---|
| Ampere | RTX A6000, A100 | BF16, or a 4-bit weight-only checkpoint (AWQ, GPTQ, MXFP4) when memory is short. FP8 checkpoints load but compute in 16-bit |
| Ada | L40, L40S, RTX 6000 Ada | FP8 from a per-channel checkpoint, which gets FP8 math. Block-scaled FP8 runs weight-only here |
| Hopper | H100 PCIe, H200 NVL | FP8, per-channel or block-scaled |
| Blackwell | RTX PRO 6000 on demand, B200 and B300 built to order | FP8, or NVFP4 when the model does not fit in FP8 |
Two vLLM flags cover most of this. --quantization fp8 quantizes a BF16 model when it loads, every linear layer except the output layer, with no calibration data. vLLM's docs note that the latency gain is limited in this mode, so for production I would pick a pre-quantized per-channel checkpoint. --kv-cache-dtype fp8 halves the KV cache on any GPU. Deploying vLLM with Docker shows where both go, and vLLM quantization covers the quantized checkpoint formats.
Precision questions#
Is BF16 better than FP16?
For training, yes on any GPU that has it. BF16 has FP32's range, so it needs no loss scaling, and every GPU QuantaCloud rents supports it. FP16 has three more bits of precision, which helps only when values stay within its range, as in a model that was trained in FP16.
Can I use FP8 on an A100 or RTX A6000?
You can load FP8 checkpoints, but the math runs in 16-bit. Ampere has no FP8 Tensor Cores, so vLLM stores the weights in FP8 and computes weight-only through its Marlin kernels. You get half the weight memory without the FP8 speed. For FP8 math, rent an L40S, an H100 PCIe or an H200 NVL.
What is TF32?
TF32 is a Tensor Core mode for FP32 inputs, not a storage format. It rounds the inputs to 10 mantissa bits, keeps FP32's 8-bit exponent and accumulates in FP32. Ampere and every later generation support it, and PyTorch leaves it off for matrix multiplies unless you turn it on.
FP8 or INT8?
On Ada and Hopper, FP8. The L40S lists the same 733 dense TOPS for INT8 as for FP8, and per-channel FP8 checkpoints need no calibration data, while vLLM's INT8 recipe calibrates on 512 samples. INT8 is the 8-bit path with real Tensor Core math on Ampere, where the A100 lists 624 dense INT8 TOPS against 312 for BF16. vLLM does not support INT8 on compute capability 10.0 and newer, so on Blackwell use FP8.
Does FP8 lose accuracy?
Very little for most LLMs. Red Hat's per-channel FP8 build of Llama 3.3 70B recovers 100.6% of the BF16 average on the OpenLLM v1 tasks and 99.8% on v2. Check your own prompts anyway, especially math and code, and check again if you also move the KV cache to FP8.
My default is simple: BF16 for training, FP8 for serving on any GPU from Ada up, and a 4-bit format only when the model does not fit in FP8. On Ampere, serve BF16 or 4-bit weight-only and treat FP8 checkpoints as a memory saving. If FP8 math decides the budget, compare the L40S, H100 PCIe and H200 NVL, and look at the RTX PRO 6000 Blackwell when you also want FP4.
Launch an L40S for FP8