What is FP4 (NVFP4 and MXFP4)?

Abbreviated FP4

FP4 is a 4-bit floating-point number format, usually laid out as E2M1: 1 sign bit, 2 exponent bits and 1 mantissa bit. At 0.5 bytes per value it stores a model's weights in a quarter of the space BF16 needs, and Blackwell-generation GPUs can run matrix math on it directly.

Why 4 bits needs block scaling

E2M1 has only 16 bit patterns. Ignoring sign, that is 8 magnitudes: 0, 0.5, 1, 1.5, 2, 3, 4 and 6. Real model weights are not spread neatly over 0 to 6, so FP4 is never used on a raw tensor. Instead the values are split into small blocks, each block stores one shared scale factor that stretches the 8-magnitude grid to fit that block's actual size, and every element is kept as a 4-bit code. At use time the code is multiplied back by its block's scale.

Smaller blocks and a finer-grained scale track outliers more closely, at the cost of storing more scales. The two formats you will see make different choices:

FormatBlock sizeBlock scaleBits per value, scale included
MXFP432 values8-bit, power of two (E8M0)4.25 (4 + 8/32)
NVFP416 valuesFP8 E4M3 plus a per-tensor scale (FP32)about 4.5 (4 + 8/16)

MXFP4 comes from the Open Compute Project's Microscaling (MX) specification, an open standard rather than a single vendor's format. Its scale can only be a power of two. NVFP4 is NVIDIA's format for Blackwell: blocks are half the size, and the FP8 block scale can take non-power-of-two values, with a second per-tensor scale on top. Some open models ship with MXFP4 weights, for example OpenAI's gpt-oss.

Which GPUs run FP4 natively

Native FP4 Tensor Core math is a Blackwell-generation feature: the B200, the B300 and the RTX 50-series cards. The H100 has no native FP4 support.

Older GPUs can still store 4-bit weights. They just cannot run FP4 math: the weights are unpacked to a 16-bit type just before each multiply, the same way existing weight-only 4-bit schemes work. That keeps the memory saving, and memory use is what decides whether a model fits, but it does not give the speedup native FP4 hardware gets.

What it means when you pick a GPU

FP4 weights cost 0.5 bytes per parameter before scales. Llama 3.1 70B has 70 billion parameters: 70 billion x 0.5 bytes = 35 GB, against 140 GB in BF16. Counting the block scales, MXFP4 comes to 70 billion x 4.25 / 8 = 37.2 GB and NVFP4 to 70 billion x 4.5 / 8 = 39.4 GB. In practice some tensors are often kept at higher precision, so budget a little more.

That changes which cards work. A 70B model at 4 bits fits on one H100 with 80 GB of HBM3, leaving roughly 40 GB for the KV cache. It does not fit on a 32 GB RTX 5090 even at 4 bits, because the weights alone are over 35 GB. Llama 3.1 405B has 405 billion parameters, so NVFP4 weights come to 405 billion x 4.5 / 8 = 228 GB: too big for a B200 (180 GB HBM3e), but within the 288 GB of HBM3e on a B300 with about 60 GB left over.

Two cautions. These are weights only, so add the KV cache, activations and framework overhead. And an FP4 checkpoint on an H100 or A100 buys you the memory saving, not native FP4 speed, so if throughput is the goal, look at the Blackwell cards.

To check a specific model and precision against a card before you rent, use the VRAM calculator. To see how FP4 compares with other precisions that cut memory, read the quantization entry.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.