NVFP4 is NVIDIA's 4-bit floating-point format for Blackwell GPUs. Each value is a 4-bit float with 1 sign bit, 2 exponent bits and 1 mantissa bit, so its magnitude can only be 0, 0.5, 1, 1.5, 2, 3, 4 or 6. Every block of 16 values shares an FP8 scale, and each tensor gets one more scale in FP32, which works out to 4.5 bits per value. Blackwell Tensor Cores compute on it directly. Older NVIDIA GPUs can still load an NVFP4 checkpoint in vLLM, but they unpack the weights and do the math in 16-bit.
The reason to care is memory. Llama 3.3 70B takes 141.1 GB in BF16, 72.7 GB in FP8 and 42.7 GB in NVFP4, so on one 96 GB card it leaves room for long conversations instead of short ones. On QuantaCloud, the GPU that computes NVFP4 natively on demand is the RTX PRO 6000 Blackwell, from the console price, and B200 and B300 servers are built to order.
Launch an RTX PRO 6000 BlackwellHow NVFP4 stores a number#
NVFP4 stores every value in two parts: a 4-bit number and the scales that stretch it back to size. NVIDIA's Transformer Engine documentation writes it as x = x_e2m1 * s_block * s_global:
| Part | Format | Shared by |
|---|---|---|
| The value | E2M1: 1 sign, 2 exponent and 1 mantissa bit, magnitudes up to 6 | One value |
| The block scale | FP8 E4M3 | 16 consecutive values |
| The tensor scale | FP32 | The whole tensor |
The block scale is set from the largest magnitude in its 16 values divided by 6, so the biggest value in each block lands at or near 6, the top of the 4-bit grid, and the others fall into place below it. The FP32 tensor scale keeps those block scales inside the range FP8 can store, whose maximum is 448. The overhead is one 8-bit scale per 16 values, which is where NVIDIA's 4.5 bits per value comes from (our calculation: 4 + 8 / 16 = 4.5).
NVFP4 vs MXFP4 vs FP8#
The difference between the two 4-bit formats is the scale, and the difference from FP8 is the number of bits:
| NVFP4 | MXFP4 | FP8 (E4M3) | |
|---|---|---|---|
| Value | 4-bit E2M1, up to 6 | 4-bit E2M1, up to 6 | 8-bit E4M3, up to 448 |
| Values per scale | 16 | 32 | A whole tensor, a channel or a 128 x 128 block, depending on the checkpoint |
| Scale format | FP8 E4M3, plus FP32 per tensor | E8M0, a power of two | A higher-precision float, FP32 in vLLM's own FP8 modes |
| Bits per value, scales included | 4.5 | 4.25 | About 8 |
| Tensor Core math | Blackwell | Blackwell | Ada, Hopper, Blackwell |
| Where you meet it | NVIDIA's and Red Hat's NVFP4 checkpoints | gpt-oss-120b and gpt-oss-20b | FP8 checkpoints such as Qwen3-32B-FP8 |
MXFP4 is the Open Compute Project's microscaling format, and NVIDIA calls it NVFP4's predecessor. NVFP4 halves the block from 32 values to 16 and replaces the power-of-two scale with an FP8 one that can take fractional values, and NVIDIA says both changes reduce quantization error. MXFP4 is what OpenAI shipped gpt-oss in: the 120b checkpoint is 65.2 GB, and vLLM's MXFP4 path needs compute capability 8.0 or newer, which every GPU QuantaCloud rents meets (gpt-oss GPU requirements).
Against FP8, NVFP4 halves the bits again. NVIDIA puts the saving at about 1.8x against FP8 and 3.5x against FP16. When it quantized DeepSeek-R1-0528 from its original FP8 to NVFP4, NVIDIA measured 1% or less accuracy loss on its key language tasks, and 2% better on AIME 2024. The FP8 vs FP16 vs BF16 guide covers the 8-bit and 16-bit formats in the same detail.
Which GPUs compute NVFP4#
FP4 math is Blackwell only. In NVIDIA's PTX instruction set, 4-bit E2M1 matrix instructions first appear with the Blackwell targets: the sm_100 family (B200, B300) and sm_120 (RTX PRO 6000 Blackwell and the RTX 50 series). FP8 matrix math starts at sm_89, the Ada generation.
| GPU | Compute capability | FP8 math | NVFP4 math | On QuantaCloud |
|---|---|---|---|---|
| RTX A6000 | 8.6 | No | No | On demand |
| A100 80GB | 8.0 | No | No | On demand |
| L40, L40S, RTX 6000 Ada | 8.9 | Yes | No | On demand |
| H100 PCIe, H200 NVL | 9.0 | Yes | No | On demand |
| RTX PRO 6000 Blackwell | 12.0 | Yes | Yes | On demand |
| B200 | 10.0 | Yes | Yes | Built to order |
| B300 | 10.3 | Yes | Yes | Built to order |
| GB200 NVL72, GB300 NVL72 | 10.0, 10.3 | Yes | Yes | Built to order |
On paper, FP4 at least doubles FP8 on the same card. NVIDIA lists the B200 at 9 dense FP4 PFLOPS against 4.5 for FP8, the B300 at 14 dense FP4 PFLOPS against the same 4.5, and both at 18 FP4 PFLOPS with sparsity.
For the RTX PRO 6000 Blackwell, NVIDIA's RTX PRO Blackwell architecture whitepaper gives the Workstation Edition 2,015.2 dense FP4 TFLOPS against 1,007.6 for FP8, and twice both with sparsity. The 4 FP4 and 2 FP8 PFLOPS on NVIDIA's Server Edition page match those sparse rates. QuantaCloud's catalog lists the card only as RTX PRO 6000 Blackwell, and nvidia-smi -q on the VM shows the edition.
What happens on GPUs without FP4#
vLLM still runs an NVFP4 checkpoint on Ampere, Ada and Hopper GPUs, as weight-only 4-bit. In vLLM 0.30.0 the native NVFP4 kernels need compute capability 10.x or 12.x and a CUDA 12.8 or newer build. On any other GPU from compute capability 7.5 up, vLLM loads the weights through its Marlin kernel, expands them to 16-bit inside each matrix multiply, and logs this warning:
Your GPU does not have native support for FP4 computation but FP4 quantization is being used. Weight-only FP4 compression will be used leveraging the Marlin kernel. This may degrade performance for compute-heavy workloads.
So an H100 or A100 keeps the memory saving and loses the FP4 math. NVIDIA's own inference guide explains where that hurts: reading the prompt (prefill) "effectively saturates GPU utilization", while generating tokens one at a time is "a memory-bound operation". Smaller weights still help the second part, and the first is where vLLM's warning about compute-heavy workloads applies. vLLM logs the kernel it chose at startup, in a line of the form Using ... for NVFP4 GEMM, so you can check it on any card:
docker logs vllm 2>&1 | grep -E "NVFP4 GEMM|native support for FP4"
The memory saving is often worth it anyway. This is what Llama 3.3 70B leaves for the KV cache on each card at vLLM's default of 92% of the memory the card reports in nvidia-smi, with ECC on for the 48 GB cards (our calculation, 320 KiB of BF16 KV per token, before activations):
| GPU memory | NVFP4 (42.7 GB) leaves | FP8 (72.7 GB) leaves | NVFP4 path in vLLM |
|---|---|---|---|
| 48 GB: RTX A6000, L40S | 1.6 GiB, about 5,300 tokens | Does not fit | Weight-only (Marlin) |
| 80 GB: A100 | 33.8 GiB, about 111,000 tokens | 5.9 GiB, about 19,000 tokens | Weight-only (Marlin) |
| 80 GB: H100 PCIe | 33.5 GiB, about 110,000 tokens | 5.6 GiB, about 18,000 tokens | Weight-only (Marlin) |
| 96 GB: RTX PRO 6000 Blackwell | 48.2 GiB, about 158,000 tokens | 20.3 GiB, about 66,000 tokens | Native FP4 |
| 141 GB: H200 NVL | 89.4 GiB, about 293,000 tokens | 61.5 GiB, about 200,000 tokens | Weight-only (Marlin) |
Mistral treats NVFP4 the same way: it publishes Mistral Large 3 (675B) in NVFP4, 403.1 GB of files against 681.5 GB for the FP8 release, and says the NVFP4 build runs on a single node of H100s or A100s. Image models follow the same rule in ComfyUI, which computes in NVFP4 only on compute capability 10 and newer and needs a cu130 PyTorch build to do it. Without that build, its developers say NVFP4 can be up to 2x slower than fp8.
How much memory NVFP4 saves#
Real checkpoints save less than NVIDIA's 3.5x, and smaller models save the least. These are the file sizes on the Hugging Face API on 2026-09-28:
| Model | BF16 | FP8 | NVFP4 | NVFP4 against BF16 | NVFP4 against FP8 |
|---|---|---|---|---|---|
| Llama 3.3 70B Instruct | 141.1 GB | 72.7 GB | 42.7 GB | 3.3x smaller | 1.7x smaller |
| Qwen3-32B | 65.5 GB | 34.3 GB | 20.7 GB | 3.2x | 1.7x |
| Qwen3-8B | 16.4 GB | 9.4 GB | 6.4 GB | 2.6x | 1.5x |
| DeepSeek-R1-0528 | Released in FP8 | 688.6 GB | 423.6 GB | Not applicable | 1.6x |
| FLUX.1 [dev], image model | 23.8 GB | 12.3 GB | 9.2 GB | 2.6x | 1.3x |
The BF16 rows are the original releases. The FP8 and NVFP4 rows are Red Hat's builds, except DeepSeek, which released FP8 itself and whose NVFP4 build is NVIDIA's. The FLUX files are all Black Forest Labs' own.
Two things eat into the saving. The first is layers that stay in BF16, usually the embeddings and the output layer: Qwen3-8B keeps 1.24 billion parameters in BF16, 2.5 GB of its 6.4 GB. The second is the scales. The 42.7 GB NVFP4 Llama holds 34.2 GB of packed 4-bit weights, 4.3 GB of FP8 block scales and 4.2 GB of BF16 layers (our calculation from the tensor counts on Hugging Face). Use the real file size when you size a GPU, and add the KV cache on top, as the KV cache guide shows.
What NVFP4 costs in accuracy#
The honest answer is a little, and more on math than on multiple choice. NVIDIA's model card for its NVFP4 Llama 3.3 70B lists these scores:
| Benchmark | BF16 | NVFP4 |
|---|---|---|
| MMLU | 83.3 | 81.1 |
| GSM8K, chain of thought | 95.3 | 92.6 |
| ARC Challenge | 93.7 | 93.3 |
| IFEval | 92.1 | 92.0 |
Red Hat's model cards show the same pattern in their own runs. Its NVFP4 build of the same model recovers 98.5% of the BF16 average on the OpenLLM v1 tasks, but 90.5% on GSM8K and 94.0% on MMLU-Pro. Its FP8 build recovers 100.6% and 99.8% of the averages on the v1 and v2 task sets. The two cards come from separate evaluation runs, so compare the recoveries, not the raw scores.
The rule I follow is to run my own prompts through the FP8 and NVFP4 checkpoints of a model before switching anything to NVFP4, and to look hardest at math, code and long reasoning.
Run an NVFP4 model with vLLM#
The same vllm serve command runs NVFP4, because vLLM reads the quantization method from the checkpoint's config. On an RTX PRO 6000 with the Bare Metal template, check the GPU and driver first:
nvidia-smi --query-gpu=name,driver_version,compute_cap --format=csv
The compute capability should read 12.0. The vllm/vllm-openai:v0.30.0 image needs driver 580 or newer, and v0.30.0-cu129 is the build for older drivers. Docker also has to reach the GPU: running Docker with a GPU checks that with one command and installs the NVIDIA Container Toolkit if it is missing. Then start the server with Red Hat's NVFP4 Llama 3.3 70B:
export VLLM_API_KEY=$(openssl rand -hex 32)
docker run -d --name vllm --gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:v0.30.0 \
RedHatAI/Llama-3.3-70B-Instruct-NVFP4 \
--max-model-len 32768 \
--api-key "$VLLM_API_KEY"
When the log shows Application startup complete, read the three lines that matter:
docker logs vllm 2>&1 | grep -E "NVFP4 GEMM|native support for FP4|GPU KV cache size"
On Blackwell there is no Marlin warning, and the KV cache line tells you how many tokens fit next to the weights. Red Hat's repo is not gated, so the download needs no Hugging Face token, but Meta's Llama 3.3 license still applies to the build. The API key, Compose file and SSH tunnel are in the vLLM Docker guide, and vLLM quantization covers the other quantized formats vLLM serves.
Where to run NVFP4 on QuantaCloud#
On demand, the RTX PRO 6000 Blackwell was the only GPU in QuantaCloud's catalog with FP4 Tensor Cores on 2026-09-27. It has 96 GB of GDDR7 and compute capability 12.0, and comes 1, 2 or 4 to a VM from the console price. The cards have no NVLink, so to split one model across them, vLLM's docs recommend pipeline parallelism over tensor parallelism. The 1x configuration had 725 GB of disk in the catalog on 2026-09-27, plenty for a 42.7 GB checkpoint and the vLLM image. The disk is deleted when you stop the VM, so the next launch downloads both again.
For NVFP4 at scale, we build B200 and B300 servers to order, from one 8-GPU HGX server to InfiniBand clusters. Two details matter when you plan that move. Kernels compiled for the RTX PRO 6000's sm_120 do not run as-is on the B200's sm_100 or the B300's sm_103. And Transformer Engine supports NVFP4 training only on compute capability 10.0 and 10.3, the B200 and B300. B300 vs B200 compares the two, and the configuration, lead time and terms come back in writing when you send a capacity brief.
NVFP4 questions#
Is NVFP4 the same as FP4?
Not quite. FP4, or E2M1, is the 4-bit number itself. NVFP4 is NVIDIA's way of using it: E2M1 values with an FP8 scale for every 16 of them and an FP32 scale per tensor. MXFP4 uses the same E2M1 values with a power-of-two scale for every 32.
Can I run NVFP4 models on an H100 or H200?
Yes, in vLLM, as weight-only 4-bit. The checkpoint loads at its NVFP4 size and the math runs in 16-bit through the Marlin kernel, so you get the memory saving without the FP4 speed. On an 80 GB H100 PCIe, that is the difference between Llama 3.3 70B in NVFP4 with about 110,000 tokens of KV cache and the FP8 build with about 18,000, by our calculation above.
Can I train or fine-tune in NVFP4?
Training support is newer and narrower than inference. NVIDIA's Transformer Engine lists NVFP4 training on compute capability 10.0 and 10.3 only, which means the B200 and B300, and inference on 10.0 and newer. NVIDIA's own research trained a 12-billion-parameter model on 10 trillion tokens in NVFP4 and reported training loss and downstream accuracy comparable to an FP8 baseline. For a model you fine-tune yourself, the documented route is to train in BF16 or FP8 and quantize the result with NVIDIA's Model Optimizer or LLM Compressor, the two tools NVIDIA names for NVFP4.
Does gpt-oss use NVFP4?
No. gpt-oss ships in MXFP4, at 4.25 bits per value in its mixture-of-experts layers, and vLLM's MXFP4 path needs compute capability 8.0 or newer. The gpt-oss guide covers the GPUs it fits on.
Is NVFP4 better than AWQ or GPTQ 4-bit?
It is not smaller: Llama 3.3 70B is 39.8 GB as a 4-bit AWQ checkpoint and 42.7 GB in NVFP4. The difference is the math. AWQ is a weight-only method that computes in 16-bit on every GPU, while NVFP4 checkpoints such as Red Hat's also quantize the activations to FP4, so a Blackwell GPU runs the whole matrix multiply in 4-bit. Neither the NVIDIA nor the Red Hat model card compares NVFP4 with AWQ, so compare them on your own prompts.
My rule: use NVFP4 when a model does not fit in FP8 with the context you need, and run it on Blackwell when you want the speed as well as the memory. On an H100 or H200, treat NVFP4 as a memory format and FP8 as the speed format, and on an A100, which has no FP8 math, treat both as memory formats. Either way, check the quality on your own prompts first. Start with the NVFP4 Llama 3.3 70B on one RTX PRO 6000, read the KV cache line in the log, and size up from there with how much VRAM you need.
Launch an RTX PRO 6000 Blackwell for NVFP4