What is FP4 (NVFP4 and MXFP4)?
Abbreviated FP4
FP4 is a 4-bit floating-point number format, usually laid out as E2M1: 1 sign bit, 2 exponent bits and 1 mantissa bit. At 0.5 bytes per value it stores a model's weights in a quarter of the space BF16 needs, and Blackwell-generation GPUs can run matrix math on it directly.
Why 4 bits needs block scaling
E2M1 has only 16 bit patterns. Ignoring sign, that is 8 magnitudes: 0, 0.5, 1, 1.5, 2, 3, 4 and 6. Real model weights are not spread neatly over 0 to 6, so FP4 is never used on a raw tensor. Instead the values are split into small blocks, each block stores one shared scale factor that stretches the 8-magnitude grid to fit that block's actual size, and every element is kept as a 4-bit code. At use time the code is multiplied back by its block's scale.
Smaller blocks and a finer-grained scale track outliers more closely, at the cost of storing more scales. The two formats you will see make different choices:
| Format | Block size | Block scale | Bits per value, scale included |
|---|---|---|---|
| MXFP4 | 32 values | 8-bit, power of two (E8M0) | 4.25 (4 + 8/32) |
| NVFP4 | 16 values | FP8 E4M3 plus a per-tensor scale (FP32) | about 4.5 (4 + 8/16) |
MXFP4 comes from the Open Compute Project's Microscaling (MX) specification, an open standard rather than a single vendor's format. Its scale can only be a power of two. NVFP4 is NVIDIA's format for Blackwell: blocks are half the size, and the FP8 block scale can take non-power-of-two values, with a second per-tensor scale on top. Some open models ship with MXFP4 weights, for example OpenAI's gpt-oss.
Which GPUs run FP4 natively
Native FP4 Tensor Core math is a Blackwell-generation feature: the B200, the B300 and the RTX 50-series cards. The H100 has no native FP4 support.
Older GPUs can still store 4-bit weights. They just cannot run FP4 math: the weights are unpacked to a 16-bit type just before each multiply, the same way existing weight-only 4-bit schemes work. That keeps the memory saving, and memory use is what decides whether a model fits, but it does not give the speedup native FP4 hardware gets.
What it means when you pick a GPU
FP4 weights cost 0.5 bytes per parameter before scales. Llama 3.1 70B has 70 billion parameters: 70 billion x 0.5 bytes = 35 GB, against 140 GB in BF16. Counting the block scales, MXFP4 comes to 70 billion x 4.25 / 8 = 37.2 GB and NVFP4 to 70 billion x 4.5 / 8 = 39.4 GB. In practice some tensors are often kept at higher precision, so budget a little more.
That changes which cards work. A 70B model at 4 bits fits on one H100 with 80 GB of HBM3, leaving roughly 40 GB for the KV cache. It does not fit on a 32 GB RTX 5090 even at 4 bits, because the weights alone are over 35 GB. Llama 3.1 405B has 405 billion parameters, so NVFP4 weights come to 405 billion x 4.5 / 8 = 228 GB: too big for a B200 (180 GB HBM3e), but within the 288 GB of HBM3e on a B300 with about 60 GB left over.
Two cautions. These are weights only, so add the KV cache, activations and framework overhead. And an FP4 checkpoint on an H100 or A100 buys you the memory saving, not native FP4 speed, so if throughput is the goal, look at the Blackwell cards.
To check a specific model and precision against a card before you rent, use the VRAM calculator. To see how FP4 compares with other precisions that cut memory, read the quantization entry.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.
Tensor Core
A Tensor Core is the GPU hardware unit that executes an entire matrix multiply-accumulate as one instruction instead of one scalar multiply at a time. How that trade unlocks NVIDIA's highest FLOP counts, and why an H100 has only four of them per SM.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
TFLOPS
TFLOPS means trillions of floating-point operations per second, a GPU's peak math speed. Why precision changes the number and why it won't predict LLM speed.
QLoRA
QLoRA fine-tunes an LLM by freezing its weights in 4-bit and training small LoRA adapters on top, so a 65B model tunes on one 48GB GPU.