What is FP8?
Abbreviated FP8
FP8 is a family of 8-bit floating-point number formats for deep learning, and in practice it means storing a model's weights (and often its activations) in one byte each instead of the two bytes that BF16 uses. The formats were proposed in the paper "FP8 Formats for Deep Learning" (Micikevicius et al., submitted September 2022), which defines two variants: E4M3, with a 4-bit exponent and a 3-bit mantissa, and E5M2, with a 5-bit exponent and a 2-bit mantissa.
Why two variants
Both variants spend one byte on a sign, an exponent and a mantissa, but they split the bits differently. E4M3 gives more precision per value, because it has one more mantissa bit. E5M2 gives more dynamic range, because it has one more exponent bit. The paper presents them as a pair so that different tensors can use the one that suits them. The authors report that FP8 matched 16-bit training quality across convolutional, recurrent and transformer models, including language models up to 175 billion parameters, and that FP8 post-training quantization worked for models that had resisted fixed-point INT8 quantization.
FP8 versus INT8 and BF16
FP8 is a floating-point format, so each value carries its own exponent, unlike INT8 fixed-point values. The paper reports that FP8 post-training quantization worked for models that resisted fixed-point INT8 quantization. Against BF16, FP8 halves the bytes per value, which halves the weight footprint and the memory traffic for every matrix multiply. It is one of the formats under the broader heading of quantization, and sits between BF16 and the 4-bit formats such as FP4.
Hardware matters here. The speedup quoted on the Llama card below was measured on H100 GPUs, whose tensor cores are the units that run low-precision matrix math. Check the card and your runtime for which GPU generations a given FP8 checkpoint supports.
FP8 models in the catalog
Most FP8 repos on Hugging Face are re-uploads of an existing model with its linear layers quantized, and the card says what was quantized. Examples:
- Llama 3.1 8B Instruct FP8 from NVIDIA. Its card says the weights and activations of the linear operators inside transformer blocks are quantized to FP8 with TensorRT Model Optimizer, which reduces disk size and GPU memory by approximately 50%. The card reports an approximately 1.3x speedup on H100 GPUs.
- Llama 3.1 70B Instruct FP8, the same treatment for a larger model.
- GLM 4.5 Air FP8. Its card lists 106 billion total parameters with 12 billion active, an MIT license and a 128K context window, and says the FP8 version needs H100 x 4 or H200 x 2 for full context.
- Llama 3.1 405B FP8, where halving the footprint is the difference between a large multi-GPU node and a larger one.
- Hermes 3 Llama 3.1 405B FP8, a fine-tune of the same base in FP8.
Note the variety in the names. Some repos add suffixes such as "dynamic" or "block", which describe how the scaling factors are computed, so read the card for the scheme before comparing two FP8 repos.
What FP8 does not shrink
FP8 weights do not automatically mean an FP8 KV cache. The GLM card above says weights and cache are both in FP8 for that release, but for other models the cache may stay in BF16 and grow with context length and batch size. Check the card, and treat the memory number as weights plus cache plus runtime overhead.
What it means when you pick a GPU
If a model has an FP8 build, it roughly halves the weight memory compared with its BF16 original, which can move a model down one tier of VRAM. To get a speedup as well, pick a GPU generation your runtime supports for FP8 (the Llama card's speedup was measured on H100). Use the VRAM calculator to size weights and cache for your context length, and Aquanode rents GPUs by the hour so you can try the FP8 build and the BF16 build on the same model before committing.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.
FP4 (NVFP4 and MXFP4)
FP4 is a 4-bit floating-point format that stores model weights at 0.5 bytes each. NVFP4 and MXFP4 add block scaling, and native FP4 math needs Blackwell GPUs.
Tensor Core
A Tensor Core is the GPU hardware unit that executes an entire matrix multiply-accumulate as one instruction instead of one scalar multiply at a time. How that trade unlocks NVIDIA's highest FLOP counts, and why an H100 has only four of them per SM.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.