NVFP4 and MXFP4 are both 4-bit floating-point formats built on the same E2M1 element, but they scale the numbers differently. MXFP4 is the open Open Compute Project standard: blocks of 32 values share a power-of-two scale. NVFP4 is NVIDIA's format: blocks of 16 values share a finer FP8 scale, with a second per-tensor scale on top, which NVIDIA says tracks the data better at a cost of about 4.5 bits per value instead of 4.25.
This guide covers:
- How each format is built, bit by bit
- Why block size and scale format change accuracy
- The accuracy and memory numbers NVIDIA has published
- Which hardware runs which format
- How to choose, and what to check before you quantize a model
This is part of our datacenter GPU guide.
TL;DR
- Same element, different scaling. Both store each number as a 4-bit E2M1 value (1 sign, 2 exponent, 1 mantissa bit). The difference is the block of values that shares a scale factor and what format that scale is in.
- NVFP4: 16 values per block, FP8 (E4M3) scale, plus a per-tensor FP32 scale. About 4.5 bits per value, per NVIDIA.
- MXFP4: 32 values per block, E8M0 (power-of-two) scale, no second-level scale. About 4.25 bits per value (4 + 8 / 32, computed).
- Hardware: NVIDIA documents NVFP4 as implemented in Blackwell's fifth-generation tensor cores. AMD documents MXFP4 on its MI355X.
- Verdict: on Blackwell, NVFP4 is the format NVIDIA builds its published FP4 results around. MXFP4 is the portable choice if you need one checkpoint format across vendors.
What both formats share
An E2M1 number has one sign bit, two exponent bits and one mantissa bit. NVIDIA's NVFP4 post lists the values it can represent: 0, plus or minus 0.5, 1, 1.5, 2, 3, 4 and 6, roughly a range of minus 6 to 6. That is a tiny set. A raw 4-bit number cannot cover the spread of values in a neural network layer, so both formats use microscaling: split a tensor into small blocks, store one scale per block, and multiply each 4-bit value by its block's scale to reconstruct the real number.
Everything that matters comes down to two choices: how many values share a scale, and what number format the scale itself uses. The FP4 glossary entry covers the basics, and quantization covers how weights get converted.
MXFP4: the open standard
The Microscaling Formats (MX) Specification v1.0 was released on October 17, 2023 by the MX Alliance of AMD, Arm, Intel, Meta, Microsoft, NVIDIA and Qualcomm through the Open Compute Project, in an open, license-free format. It defines four formats: MXFP8, MXFP6, MXFP4 and MXINT8.
AMD's ROCm documentation lists the MXFP4 layout: FP4 (E2M1) elements, a block size of 32, and an E8M0 scale. E8M0 is eight exponent bits and no mantissa, so a scale can only be a power of two. NVIDIA's research blog on 4-bit models describes the consequence: the ideal scale often falls between two powers of two, so the tiny 4-bit budget "pays for the error."
NVFP4: NVIDIA's format
NVIDIA introduced NVFP4 on its developer blog on June 24, 2025. Its structure, from that post:
- Element: FP4 E2M1.
- Block size: 16 values, half of MXFP4's 32.
- Block scale: FP8 in the E4M3 format, which has a mantissa, so a scale can land between powers of two.
- Second-level scale: one FP32 scalar per tensor.
- Storage cost: about 4.5 bits per value, from one 4-bit value plus one FP8 scale per 16 values.
NVFP4 is not one of the four formats in the OCP MX specification. It is NVIDIA's own variant.
Why the differences matter: a worked example
This is an illustration we computed, not a benchmark. Imagine a block whose largest value is 5.0. E2M1 can represent 4 and 6, but not 5.
- With a power-of-two scale of 1, the value 5.0 has to round to 4 or 6, an error of 1.0 on the block's biggest number.
- With a finer E4M3 scale of 5 divided by 6, about 0.833, the value 5.0 scales to exactly 6 and reconstructs without error.
Real blocks have many values, so errors do not vanish this neatly, but the direction holds. A scale that can take non-power-of-two values fits the block's range more tightly. Shrinking the block from 32 to 16 values helps for the same reason: NVIDIA says the tighter grouping gives twice as many chances to match the local dynamic range of the data, which in our reading also means one outlier value distorts 16 neighbours instead of 32.
The price is overhead. NVFP4 carries one 8-bit scale per 16 values (plus a negligible per-tensor scalar), MXFP4 carries one 8-bit scale per 32.
| Property | MXFP4 | NVFP4 |
|---|---|---|
| Element format | FP4 E2M1 | FP4 E2M1 |
| Values per block | 32 | 16 |
| Block scale format | E8M0 (power of two) | FP8 E4M3 |
| Second-level scale | None | FP32 per tensor |
| Bits per value, with scales | About 4.25 (computed) | About 4.5 (NVIDIA) |
| Defined by | OCP MX Specification v1.0 | NVIDIA |
| Hardware documented | AMD CDNA 4 (MI355X) | NVIDIA Blackwell tensor cores |
What NVIDIA has published on accuracy
Nothing here is our measurement. These are NVIDIA's numbers, from its own models and settings.
Memory. NVIDIA says NVFP4 is approximately 3.5x smaller than FP16 and approximately 1.8x smaller than FP8.
Post-training quantization of DeepSeek-R1-0528. NVIDIA's NVFP4 post compares FP8 with NVFP4 on reasoning benchmarks:
| Benchmark | FP8 | NVFP4 |
|---|---|---|
| MMLU-PRO | 85% | 84% |
| GPQA Diamond | 81% | 80% |
| HLE | 15% | 14% |
| LiveCodeBench | 77% | 76% |
| SciCode | 40% | 40% |
| Math-500 | 98% | 98% |
| AIME 2024 | 89% | 91% |
NVIDIA's reading: one percentage point or less of degradation on key tasks, and 2 percentage points better on AIME 2024. One model and one vendor-run test suite is a data point, not a guarantee for your model.
Pretraining. NVIDIA's paper "Pretraining Large Language Models with NVFP4" (arXiv 2509.25149, submitted September 29, 2025, revised March 4, 2026) trained a 12-billion-parameter model on 10 trillion tokens, which the authors describe as the longest publicly documented 4-bit training run. They report training loss and downstream accuracy comparable to an FP8 baseline. The recipe is not a simple cast: it combines Random Hadamard transforms, two-dimensional quantization, stochastic rounding for gradients and keeping selected layers in higher precision.
What we could not find. We did not find an NVIDIA-published, like-for-like accuracy table of NVFP4 against MXFP4 on the same model that we could cite here. NVIDIA's argument for NVFP4's accuracy advantage in its blog is structural (the finer block and scale), supported by its own FP8 comparisons above. If the exact gap matters to you, test both on your model.
Which hardware runs which format
NVIDIA Blackwell. NVIDIA's NVFP4 post states that Blackwell's fifth-generation tensor cores implement NVFP4, handling element grouping, dynamic scaling and 4-bit matrix operations. That covers the B200 and the B300. NVIDIA lists dense FP4 at 72 PFLOPS for an eight-GPU HGX B200 board and 108 PFLOPS for an HGX B300 board, with sparse figures at 144 PFLOPS for both. Per GPU, NVIDIA's Blackwell Ultra blog gives 15 PFLOPS dense NVFP4 for B300 against 10 for Blackwell. NVIDIA's MLPerf Inference v5.1 write-up says it used NVFP4 extensively across its Blackwell and Blackwell Ultra DeepSeek-R1 and Llama submissions, with the KV cache for DeepSeek-R1 in FP8.
AMD Instinct MI355X. AMD documents hardware support for the OCP MX formats MXFP8, MXFP6 and MXFP4 in its CDNA 4 matrix cores. AMD's product page lists peak MXFP4 at 10.1 PFLOPS, with 288 GB of HBM3E and 8 TB/s of memory bandwidth. Those are AMD's peak theoretical figures. See our MI355X guide.
Hopper (H100, H200). NVIDIA's Hopper generation predates both FP4 formats. Hopper's low-precision path is FP8, covered in our Transformer Engine and FP8 guide and the FP8 glossary entry. NVIDIA's Blackwell page presents FP4 as a Blackwell feature; we did not find NVIDIA documenting native FP4 tensor math on Hopper.
How to choose
Choose NVFP4 when:
- You run on Blackwell and want the format NVIDIA's published FP4 results are built around.
- Accuracy at 4 bits matters, such as reasoning models where small errors compound over long outputs.
- You have, or can build, an NVFP4 checkpoint and an engine that runs it.
Choose MXFP4 when:
- You need a checkpoint format the open standard defines and that AMD hardware supports natively.
- You run mixed fleets and want one 4-bit format across vendors.
- The extra memory for NVFP4's finer scales is not worth it to you, or your model already ships in MXFP4.
Stay at FP8 when your accuracy budget is tight or your engine lacks FP4 kernels. FP8 on Blackwell has the same dense throughput on B200 and B300 per NVIDIA's HGX table, so you are not forced to FP4 to use these GPUs.
A short checklist before quantizing any model:
- Confirm your inference engine supports the format on your GPU. Hardware support and software support are separate.
- Calibrate on data that looks like your production traffic.
- Run your own evaluation set, with long outputs if you serve reasoning models, and compare to FP8 or BF16.
- Decide what to keep in higher precision. NVIDIA's pretraining recipe keeps selected layers at higher precision.
Cost: what 4-bit changes
The memory saving is the direct effect: roughly 3.5x smaller than FP16 and 1.8x smaller than FP8, per NVIDIA. For inference that means fewer GPUs per replica and more room for KV cache. The throughput effect is workload-specific, so we do not quote a tokens-per-hour figure for FP4 in general. The closest published anchor is NVIDIA's MLPerf Inference v5.1 DeepSeek-R1 result on GB300 NVL72 (NVFP4 weights): 5,842 tokens per second per GPU offline, which is about 21.0 million tokens per GPU-hour (computed from NVIDIA's per-GPU division of system throughput). That is a tuned rack-scale benchmark, so expect less. Multiply your own measured tokens per GPU-hour by the live hourly price in the box below, and compare against the same math at FP8. See /pricing and the GPU index for current prices across GPUs.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Both are Blackwell parts with NVFP4 support, and the box shows live prices or "None right now."
For a head-to-head of the two chips, read B300 vs B200.
What's next
NVIDIA's Rubin page says the Vera Rubin NVL72 platform has a "new Transformer Engine with adaptive compression to boost NVFP4 inference performance," with no NVFP4 performance figures given. Read the Rubin guide for what is and is not confirmed. AMD's next generation is covered in the MI400 and MI450 guide.
FAQ
What is the difference between NVFP4 and MXFP4?
Both use 4-bit E2M1 values. MXFP4 shares an E8M0 power-of-two scale across 32 values. NVFP4 shares an FP8 E4M3 scale across 16 values and adds a per-tensor FP32 scale.
Is NVFP4 more accurate than MXFP4?
NVIDIA argues yes, because the smaller block and fractional scale match local data ranges more closely. It publishes NVFP4 against FP8 comparisons, not a direct MXFP4 table we could cite. Test on your own model.
Which GPUs support NVFP4?
NVIDIA documents NVFP4 in Blackwell's fifth-generation tensor cores, which means B200, B300 and the GB200 and GB300 systems. We found no NVIDIA documentation of native FP4 on Hopper.
Which GPUs support MXFP4?
AMD documents native MXFP4 on the Instinct MI355X through CDNA 4 matrix cores. MXFP4 is defined by the OCP MX specification, which NVIDIA co-authored.
How many bits per value does each format use?
NVIDIA states about 4.5 bits per value for NVFP4. For MXFP4 the same arithmetic gives 4.25: four bits plus an 8-bit scale shared by 32 values.
Can I train in FP4?
NVIDIA has published a pretraining run of a 12B-parameter model on 10 trillion tokens in NVFP4 with results comparable to FP8, using a specific recipe. It is not a drop-in switch.
Sources
- NVIDIA, Introducing NVFP4 for efficient and accurate low-precision inference (June 24, 2025): https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference
- NVIDIA Research, Pushing intelligence to 4-bit (June 29, 2026): https://research.nvidia.com/labs/eai/blogs/pushing-intelligence-to-4-bit/
- NVIDIA, Pretraining Large Language Models with NVFP4 (arXiv 2509.25149): https://arxiv.org/abs/2509.25149
- Open Compute Project, MX Specification v1.0 announcement (October 17, 2023): http://www.opencompute.org/news/amd-arm-intel-meta-microsoft-nvidia-and-qualcomm-standardize-next-generation-narrow-precision-data-formats-for-ai/
- AMD ROCm blog, High-accuracy MXFP4, MXFP6 and mixed-precision models on AMD GPUs: https://rocm.blogs.amd.com/software-tools-optimization/mxfp4-mxfp6-quantization/README.html
- AMD Instinct MI355X product page: https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html
- NVIDIA HGX platform: https://www.nvidia.com/en-us/data-center/hgx/
- NVIDIA, Inside NVIDIA Blackwell Ultra: https://developer.nvidia.com/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era/
- NVIDIA, Blackwell Ultra sets new inference records in MLPerf debut (MLPerf Inference v5.1): https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/
- NVIDIA Rubin platform page: https://www.nvidia.com/en-us/data-center/technologies/rubin/