The rule I follow is to pick the quantization format by GPU generation, because vLLM runs some formats at full speed only on newer GPUs and falls back to a weight-only path on older ones. On the RTX A6000 and A100 (Ampere), 4-bit AWQ or GPTQ is the format to use, and FP8 only saves memory. On the L40S, L40 and RTX 6000 Ada, use per-channel FP8 checkpoints, which get FP8 math. On the H100 PCIe and H200 NVL, any FP8 checkpoint gets FP8 math. In the 2026-09-27 catalog, the RTX PRO 6000 was the only on-demand GPU that computes in NVFP4. In every case vLLM reads the format from the checkpoint, so a pre-quantized model needs no --quantization flag.
The formats vLLM 0.30 loads#
vLLM reads a checkpoint's quantization_config and picks the matching method, and the table lists the ones you will meet on Hugging Face. The --quantization value is only needed to override that detection or to quantize a 16-bit model while it loads.
| Format | What is stored | Example checkpoint | Method vLLM detects |
|---|---|---|---|
| AWQ | 4-bit integer weights with a scale per group, 16-bit activations | Qwen/Qwen3-32B-AWQ | awq |
| GPTQ | 4-bit or 8-bit integer weights, 16-bit activations | RedHatAI/Qwen3-32B-quantized.w4a16 | gptq, or compressed-tensors when made with llm-compressor |
| FP8, per-channel | FP8 weights with one scale per output channel, activations scaled per token at run time | RedHatAI/Qwen3-32B-FP8-dynamic | compressed-tensors |
| FP8, block-scaled | FP8 weights with one scale per 128 x 128 block | Qwen/Qwen3-32B-FP8 | fp8 |
| NVFP4 | 4-bit floating-point weights and activations in groups of 16 | RedHatAI/Qwen3-32B-NVFP4 | compressed-tensors, or modelopt_fp4 from NVIDIA's tool |
| MXFP4 | 4-bit floating-point weights in groups of 32, as gpt-oss ships | openai/gpt-oss-120b | mxfp4, handled as gpt_oss_mxfp4 for gpt-oss |
| INT8 W8A8 | 8-bit integer weights and activations | llm-compressor w8a8 checkpoints | compressed-tensors |
Two formats moved out of vLLM. bitsandbytes now needs vllm-bnb-plugin and GGUF needs vllm-gguf-plugin, and vLLM's docs call its GGUF support highly experimental and under-optimized. If a model only exists as GGUF, Ollama or llama.cpp is the natural server for it, and vLLM vs Ollama covers what changes when you switch.
What each format needs from the GPU#
The dividing line is compute capability, as vLLM 0.30.0's kernel code draws it. FP8 math needs 8.9 or newer, block-scaled FP8 math needs 9.0 or newer, and FP4 math needs Blackwell. When a GPU cannot run a format's math, vLLM keeps the weights compressed and runs them through its Marlin kernel as 16-bit math: you keep the memory saving and lose the speed of the low-precision Tensor Cores.
| Format | RTX A6000, A100 (Ampere, 8.6 and 8.0) | L40S, L40, RTX 6000 Ada (Ada, 8.9) | H100 PCIe, H200 NVL (Hopper, 9.0) | RTX PRO 6000 (Blackwell, 12.0) |
|---|---|---|---|---|
| AWQ, GPTQ 4-bit | 4-bit weights, 16-bit math | 4-bit weights, 16-bit math | 4-bit weights, 16-bit math | 4-bit weights, 16-bit math |
| FP8, per-channel | Weight-only | FP8 math | FP8 math | FP8 math |
| FP8, block-scaled | Weight-only | Weight-only | FP8 math | FP8 math |
| NVFP4 | Weight-only | Weight-only | Weight-only | FP4 math |
| MXFP4 | Loads, 8.0 and newer | Loads | Loads | Loads |
AWQ and GPTQ never compute in 4 bits. vLLM unpacks the weights inside its kernels, Marlin on every GPU here and Machete on Hopper when it supports the layout, so their benefit is memory and the bandwidth saved reading smaller weights. MXFP4 loads on every QuantaCloud GPU, but vLLM's gpt-oss recipe still lists Ampere and Ada as work in progress, which the gpt-oss GPU guide covers.
The weight-only fallback announces itself. vLLM logs a warning that begins Your GPU does not have native support for FP8 computation but FP8 quantization is being used, or the same sentence for FP4, and names Marlin. If you see it on a card you chose for its FP8 or FP4 Tensor Cores, you loaded the wrong checkpoint for that card. NVFP4 explained and FP8 vs FP16 vs BF16 go deeper into the number formats.
Memory saved, from real checkpoint sizes#
Quantization roughly halves the weights at 8 bits and cuts them by about 70% at 4 bits. The sizes below are the checkpoint files on Hugging Face, which run larger than the parameter count suggests because the embedding and output layers stay 16-bit and 4-bit formats carry scales.
| Checkpoint | Format | Size | Smaller than BF16 (our calculation) |
|---|---|---|---|
Qwen/Qwen3-32B | BF16 | 65.5 GB | |
RedHatAI/Qwen3-32B-FP8-dynamic | FP8, per-channel | 34.3 GB | 48% |
Qwen/Qwen3-32B-FP8 | FP8, block-scaled | 34.3 GB | 48% |
Qwen/Qwen3-32B-AWQ | AWQ, 4-bit | 19.3 GB | 71% |
RedHatAI/Qwen3-32B-quantized.w4a16 | GPTQ, 4-bit | 19.2 GB | 71% |
RedHatAI/Qwen3-32B-NVFP4 | NVFP4 | 20.7 GB | 68% |
meta-llama/Llama-3.3-70B-Instruct | BF16 | 141.1 GB | |
RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic | FP8, per-channel | 72.7 GB | 49% |
casperhansen/llama-3.3-70b-instruct-awq | AWQ, 4-bit | 39.8 GB | 72% |
nvidia/Llama-3.3-70B-Instruct-NVFP4 | NVFP4 | 42.7 GB | 70% |
NVFP4 files come out a little larger than AWQ and GPTQ, because every 16 values carry an FP8 scale. The Llama 3.3 checkpoints carry Meta's Llama 3.3 licence, and the base repo is gated. The Qwen3 checkpoints are Apache-2.0.
What fits on 48, 80, 96 and 141 GB#
The saving matters because of what it leaves for the KV cache. vLLM takes 92% of the GPU by default, and after the weights the rest holds the context of every running request: 256 KiB per token for Qwen3-32B and 320 KiB for Llama 3.3 70B at BF16. This table is our calculation from the memory each card reports in nvidia-smi, with ECC on for the 48 GB cards, before activations and CUDA graphs take their share:
| Weights (file size) | 48 GB: RTX A6000, L40S | 80 GB: A100 | 80 GB: H100 PCIe | 96 GB: RTX PRO 6000 | 141 GB: H200 NVL |
|---|---|---|---|---|---|
| Qwen3-32B BF16 (61.0 GiB) | Does not fit | 12.6 GiB, about 51,000 tokens | 12.3 GiB, about 50,000 | 26.9 GiB, about 110,000 | 68.2 GiB, about 279,000 |
| Qwen3-32B FP8 (32.0 GiB) | 9.4 GiB, about 38,000 tokens | 41.6 GiB, about 170,000 | 41.3 GiB, about 170,000 | 56.0 GiB, about 230,000 | 97.2 GiB, about 400,000 |
| Qwen3-32B AWQ (18.0 GiB) | 23.4 GiB, about 96,000 tokens | 55.6 GiB, about 228,000 | 55.3 GiB, about 226,000 | 69.9 GiB, about 286,000 | 111.2 GiB, about 455,000 |
| Llama 3.3 70B FP8 (67.7 GiB) | Does not fit | 5.9 GiB, about 19,000 tokens | 5.6 GiB, about 18,000 | 20.3 GiB, about 66,000 | 61.5 GiB, about 200,000 |
| Llama 3.3 70B AWQ (37.0 GiB) | 4.4 GiB, about 14,000 tokens | 36.6 GiB, about 120,000 | 36.2 GiB, about 119,000 | 50.9 GiB, about 167,000 | 92.1 GiB, about 302,000 |
| Llama 3.3 70B NVFP4 (39.8 GiB) | 1.6 GiB, about 5,000 tokens | 33.8 GiB, about 111,000 | 33.5 GiB, about 110,000 | 48.2 GiB, about 158,000 | 89.4 GiB, about 293,000 |
Read each cell as the pool all running requests share. A 32B model at 4 bits on a 48 GB card has more room for context than the same model at FP8 on it, and a 70B model at FP8 is effectively a single-conversation setup on an 80 GB card. The KV cache guide has the formula for other models, and fixing vLLM out-of-memory errors covers what to do when the cache does not fit.
Two notes apply to the table. The 48 GB column uses the 46,068 MiB, about 45 GiB, these cards report with ECC on, and an RTX A6000 with ECC off reports 49,140 MiB, which adds about 2.8 GiB to each of its cells (our calculation: 0.92 x 3 GiB). And NVIDIA's NVFP4 checkpoint of Llama 3.3 70B sets an FP8 KV cache in its quantization config, which vLLM 0.30.0 applies by default, so that row holds about twice the tokens shown.
Serve a quantized checkpoint#
A pre-quantized checkpoint runs with the same command as any other model. The setup, including the driver check that sets $VLLM_TAG and the API key, is in deploying vLLM with Docker. On a 48 GB Ada card such as the L40S, the per-channel FP8 build of Qwen3-32B:
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v vllm-cache:/root/.cache/vllm \
"vllm/vllm-openai:$VLLM_TAG" \
RedHatAI/Qwen3-32B-FP8-dynamic \
--max-model-len 16384 \
--reasoning-parser qwen3 \
--api-key "$VLLM_API_KEY"
docker logs -f vllm
The limit is 16,384 tokens because this model is tight on the L40S. One 32,768-token request needs 8 GiB of KV cache, and 92% of the 45 GiB the card reports leaves about 9.4 GiB after the 32 GiB of weights, before activations and CUDA graphs take their share (our calculation). --max-model-len auto finds the real ceiling if you need a longer context.
When the log shows Application startup complete, press Ctrl+C and pull out the two lines that matter: the weight-only warning, which should be absent on this card, and the KV cache size.
docker logs vllm 2>&1 | grep -E "native support|GPU KV cache size"
Swap in Qwen/Qwen3-32B-AWQ or RedHatAI/Qwen3-32B-NVFP4 and the command stays the same. On the L40S the NVFP4 build logs the FP4 version of the warning, because it runs weight-only below Blackwell.
To quantize a 16-bit checkpoint as it loads instead, pass --quantization fp8 with the BF16 model, Qwen/Qwen3-32B here. vLLM 0.30.0 materializes each layer just in time and converts its weights to per-tensor FP8. I prefer a pre-quantized checkpoint on QuantaCloud anyway, because stopping an instance deletes its disk and the next launch downloads the weights again: 34.3 GB for the FP8 checkpoint against 65.5 GB for the BF16 one.
The KV cache has its own setting. --kv-cache-dtype fp8 halves the cache on any of these GPUs, but without calibration it uses scales of 1.0, so check output quality on your own prompts. vLLM's docs recommend calibrating with llm-compressor, and --kv-cache-dtype-skip-layers sliding_window keeps sensitive sliding-window layers at full precision.
Two kinds of checkpoint fail to load in 0.30.0 for reasons that have nothing to do with the GPU. GPTQ files made with group activation ordering stop with GPTQ group activation ordering (desc_act=True) is no longer supported, so pick one made with desc_act=False or static ordering. bitsandbytes and GGUF files need their plugin installed into the image, for example with a two-line Dockerfile that starts FROM vllm/vllm-openai:v0.30.0 and runs uv pip install --system vllm-gguf-plugin.
How to choose on QuantaCloud GPUs#
The GPU you rent narrows the choice to one or two formats. Prices per GPU-hour:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | - | Not listed | No |
| H200 NVL | - | Not listed | No |
Prices checked 6 Oct 2026, 02:18 UTC
On the RTX A6000 and the A100, use 4-bit AWQ or GPTQ for anything that does not fit in 16 bits. FP8 checkpoints load on these cards, but only as a memory saving. On the L40S, L40 and RTX 6000 Ada, choose per-channel FP8 checkpoints, the FP8-dynamic builds from llm-compressor, over block-scaled ones such as Qwen's own FP8 releases, which run weight-only below Hopper. On the H100 PCIe and H200 NVL, either FP8 flavour gets FP8 math, and NVFP4 saves memory but no more than a 4-bit AWQ checkpoint does. On the RTX PRO 6000, NVFP4 is the format that uses the card's FP4 Tensor Cores, and FP8 runs natively too.
vLLM quantization FAQ#
Do I need the --quantization flag?
No, not for a pre-quantized checkpoint. vLLM reads the method from the checkpoint's quantization_config. You need the flag only to quantize a 16-bit model while it loads, for example --quantization fp8, or to override detection.
Does FP8 work on the A100 or RTX A6000?
It loads, weight-only. Both are Ampere cards without FP8 Tensor Cores, so vLLM runs FP8 weights through its Marlin kernel with 16-bit math and logs a warning that says so. You get the memory saving and not the FP8 speed.
Is NVFP4 worth it on an H100?
Only for the memory. Hopper has FP8 Tensor Cores but no FP4 ones, so vLLM runs NVFP4 weight-only there, and a 4-bit AWQ or GPTQ checkpoint of the same model is slightly smaller. On the RTX PRO 6000, NVFP4 runs in FP4.
How do I quantize my own model?
vLLM's docs point to llm-compressor, which produces FP8, INT8, INT4 and NVFP4 checkpoints that vLLM loads directly, and they mark the older AutoAWQ library as deprecated. GPTQModel is the other tool the docs cover for GPTQ files.
Can I quantize the KV cache as well?
Yes. --kv-cache-dtype fp8 stores keys and values in 8 bits and halves the cache, independent of the weight format. Without a calibrated checkpoint it uses scales of 1.0, so compare outputs before you rely on it.
My rule is to match the format to the card before anything else: 4-bit AWQ or GPTQ on Ampere, per-channel FP8 on Ada, any FP8 on Hopper and NVFP4 on the RTX PRO 6000. Then size the GPU from the checkpoint's real file size plus the KV cache you need, and check the startup log for the weight-only warning on the first run. If a 32B model at FP8 is your target, a 48 GB L40S is the place to start for moderate contexts, and the GPU catalog has the larger cards for long ones.
Launch an L40S for FP8 Launch an RTX PRO 6000 for NVFP4