What is Quantization?

Quantization is the practice of storing a neural network's numbers, usually its weights and sometimes its activations, in lower-precision formats such as INT8, FP8, INT4 or FP4 instead of 16-bit floats. In an LLM it shrinks the memory footprint in direct proportion to the bit width, so a model that needs two GPUs at 16-bit can fit on one at 8-bit or 4-bit, at the cost of some accuracy.

INT8 quantization and the other formats

Quantization maps a range of real values onto a small set of levels plus a scale factor. INT8 has 256 integer levels (from -128 to 127): each weight is stored as one byte, and multiplying it by the scale gives back an approximation of the original value. In practice, small groups of weights share a scale so that one outlier does not distort the rest. INT4 does the same with 16 levels. FP8 and FP4 are floating-point formats with a few exponent bits, so they cover a wider range of magnitudes with the same number of bits. The baseline is usually a 16-bit format such as BF16; the newest 4-bit floating-point format is covered under FP4.

Weight-only quantization compresses the weights (for example 4-bit weights with 16-bit activations) and converts them back to higher precision during the matrix multiply. It saves memory and also memory reads, and since LLM decoding is limited by how fast weights stream from memory, fewer bytes per weight often means faster generation.

Weight-and-activation quantization (for example INT8 or FP8 for both) lets the Tensor Cores run the multiply in the low-precision format itself, which can raise compute throughput. It needs hardware that supports the format, and it is more sensitive because activations contain outliers. Older cards such as the A100 have no FP8 Tensor Cores, and the H100 has no native FP4 support, while the B200 does.

How accuracy is preserved

Post-training quantization (PTQ) converts an already trained model, often using a small calibration dataset to pick scales. It is fast and needs no training. Quantization-aware training (QAT) simulates low precision during training or fine-tuning so the model learns to tolerate it. It usually holds up better at very low bit widths but costs training compute.

The accuracy cost is usually small at 8-bit, larger and more variable at 4-bit, and harder to control below that. It depends on the model, the method and the task, so evaluate on your own data. Methods often keep the most sensitive parts of a model at higher precision.

What it means when you pick a GPU

Weights memory is parameters times bytes per parameter. For a 70-billion-parameter model:

PrecisionBytes per parameterWeights
16-bit (BF16)270B x 2 = 140 GB
8-bit (INT8 or FP8)170B x 1 = 70 GB
4-bit (INT4 or FP4)0.570B x 0.5 = 35 GB

Here is what that leaves free on each card (VRAM minus weights):

GPUVRAM16-bit8-bit4-bit
H10080GBDoes not fit10 GB free45 GB free
H200141GB1 GB free, unusable71 GB free106 GB free
B200180GB40 GB free110 GB free145 GB free
AMD MI300X192GB52 GB free122 GB free157 GB free
RTX 409024GBDoes not fitDoes not fitDoes not fit

The free space is not all yours. The KV cache, activations and runtime overhead come on top of the weights, so 10 GB free on an H100 at 8-bit is tight, and 1 GB free on an H200 at 16-bit is not enough to serve anything. Real 4-bit files also run a little above 35 GB because the scales are stored too, so budget 35 to 40 GB. An RTX 4090 cannot hold this model at any of these precisions. At 4-bit, a model of up to roughly 30 billion parameters (about 15 GB of weights) is a more realistic fit for its 24GB.

Quantizing the weights does not shrink the KV cache on its own. If long contexts or many concurrent users are your real constraint, size the cache separately. Put your own model, precision and context length into the VRAM calculator to see whether a given card fits. Aquanode rents GPUs by the hour, so you can pick the card that matches the precision you serve.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.