What is BF16 (bfloat16)?
Abbreviated BF16
BF16 (bfloat16, short for "brain floating point") is a 16-bit number format that keeps FP32's 8 exponent bits but cuts the mantissa to 7 bits. It stores each value in 2 bytes and covers the same range of values as FP32, which is why it is the default precision for training large models and the format many open model checkpoints are published in.
A BF16 value is 1 sign bit, 8 exponent bits and 7 mantissa bits. The layout is the top 16 bits of an FP32 number, so converting between the two is simple. The format was developed at Google Brain, which is where the name comes from.
BF16 vs FP16
Both formats take 2 bytes per value, but they spend their 16 bits differently.
| Format | Sign bits | Exponent bits | Mantissa bits | Largest finite value |
|---|---|---|---|---|
| FP32 | 1 | 8 | 23 | about 3.4e38 |
| BF16 | 1 | 8 | 7 | about 3.4e38 |
| FP16 | 1 | 5 | 10 | 65,504 |
FP16 spends more bits on precision and BF16 spends them on range. FP16 has finer steps between neighboring numbers (10 mantissa bits, a relative step of about 0.1%, against about 0.8% for BF16's 7), but its largest value is only 65,504 and very small values underflow to zero quickly.
Training produces gradients across many orders of magnitude, so FP16 training needs loss scaling: multiply the loss by a large factor so small gradients stay representable, then undo the factor before the weight update. Too large a factor overflows to infinity or NaN, and too small lets gradients vanish, so the factor is usually adjusted dynamically during the run.
BF16 shares FP32's range, so gradients and activations stay representable with no scaling step. The cost is coarser rounding. Deep learning tolerates that noise, as long as the sums inside each matrix multiply are accumulated at higher precision.
Why training defaults to BF16
BF16 mixed-precision training is a hybrid. The forward and backward matrix multiplies run in BF16 on Tensor Cores, which accumulate their partial sums in FP32. The optimizer keeps an FP32 master copy of the weights and its own state, so tiny updates are not rounded away.
Because the range matches FP32, a recipe that works in FP32 rarely needs retuning for overflow, and you drop the loss scaler entirely. That is one fewer thing that can go wrong in a multi-day run, which is the main reason BF16 replaced FP16 as the training default on hardware that supports it.
What it means when you pick a GPU
BF16 costs 2 bytes per parameter, so weights alone need parameters x 2 bytes of GPU memory. Llama 3.1 70B has 70 billion parameters: 70 billion x 2 bytes = 140 GB. A single H100 has 80 GB of HBM3, so a BF16 70B model needs at least two H100s for the weights alone. A B200 has 180 GB of HBM3e, which holds the weights on one card with about 40 GB (180 - 140) left for the KV cache and activations. An 8B model needs 8 billion x 2 = 16 GB.
Training needs far more. A common accounting from the ZeRO paper (Rajbhandari et al., 2019) for mixed-precision training with Adam is 16 bytes per parameter: 2 for the BF16 weights, 2 for gradients, 4 for the FP32 master weights and 8 for Adam's two FP32 moments. A 7B model then needs 7 billion x 16 bytes = 112 GB before activations. That is more than one 80 GB card, so full training of even a 7B model needs the state sharded across several GPUs, or far fewer trainable parameters.
Hardware support also matters. BF16 Tensor Core support starts with the Ampere generation: the A100 and the datacenter GPUs after it (H100, H200, B200), plus RTX 30-series and newer cards. The older V100 has no BF16 support, so a BF16 recipe has to be reworked into FP16 with loss scaling before it runs there. AMD's MI300X supports BF16 too. Peak dense FP16/BF16 tensor throughput is 312 TFLOPS on an A100 and 989 TFLOPS on an H100.
If 2 bytes per parameter is too much for the card you want, quantization stores weights in 8 or 4 bits, and FP4 on Blackwell GPUs brings that down to 0.5 bytes per parameter. To check what fits before you rent, put your model and precision into the VRAM calculator. Aquanode rents GPUs by the hour.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Tensor Core
A Tensor Core is the GPU hardware unit that executes an entire matrix multiply-accumulate as one instruction instead of one scalar multiply at a time. How that trade unlocks NVIDIA's highest FLOP counts, and why an H100 has only four of them per SM.
FP4 (NVFP4 and MXFP4)
FP4 is a 4-bit floating-point format that stores model weights at 0.5 bytes each. NVFP4 and MXFP4 add block scaling, and native FP4 math needs Blackwell GPUs.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
TFLOPS
TFLOPS means trillions of floating-point operations per second, a GPU's peak math speed. Why precision changes the number and why it won't predict LLM speed.
FSDP (Fully Sharded Data Parallel)
FSDP shards a model's weights, gradients and optimizer state across GPUs and gathers each layer only when needed, so models too big for one GPU can train.