What is GPTQ?
Abbreviated GPTQ
GPTQ is a post-training quantization method that compresses the weights of a generative pre-trained transformer to a few bits each, in one shot, without retraining. It comes from GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers by Elias Frantar, Saleh Ashkboos, Torsten Hoefler and Dan Alistarh, first submitted in October 2022. The paper describes it as "a new one-shot weight quantization method based on approximate second-order information."
What the paper claims
The authors report three results, all from the paper's own experiments:
- A model with 175 billion parameters can be compressed to 3 to 4 bits per weight while keeping accuracy.
- The quantization process takes approximately four GPU hours for that model.
- The compressed 175-billion-parameter model fits inside a single GPU for generative inference, and inference runs about 3.25x faster on high-end GPUs and 4.5x faster on more affordable ones, compared with the baseline the paper uses.
"One-shot" means the weights are rounded in a single pass over a small amount of data, not by training. "Second-order information" means the method uses curvature of the error, rather than just rounding each weight to its nearest value, to decide how to adjust the remaining weights as each one is rounded. See quantization for the general trade between precision and memory.
GPTQ in practice
Like AWQ, GPTQ names an algorithm, not a single file format. Hugging Face repositories labelled GPTQ carry weights produced by it, usually with the bit width in the name (for example Int4). Related names appear too: the EXL2 format is described by its own README as based on the same optimization method as GPTQ, and some repositories use a GPTQ-compatible layout produced by other tools. The GGUF format is a separate container with its own quantization types, so a GGUF file is not a GPTQ file.
Example models on Aquanode
- Vishva007 Gemma 4 E4B it W4A16 AutoRound GPTQ is a 4-bit build of Google's Gemma 4 E4B instruction-tuned model. Its card describes INT4 weights with BF16 activations (W4A16), a group size of 128, quantization with AutoRound, and calibration on 256 samples at 2048 sequence length. Only the language-model layers are quantized to INT4; the vision tower, audio tower and projectors stay in BF16. The model is subject to the Gemma Terms of Use.
This card is a useful illustration of two points. A repository named GPTQ can be produced by a different tool than the original GPTQ code, so read the card for what was actually done. And quantizing only part of a multimodal model leaves the rest at full size, which affects the memory you need.
What GPTQ does not shrink
GPTQ lowers the weight storage. It does not shrink the KV cache, which grows with context length and the number of requests, and it does not change the activations a runtime computes in higher precision. Weights stored in 4-bit are converted when the math runs, so the savings are in memory and bandwidth rather than in the arithmetic itself.
What it means when you pick a GPU
Read the bit width and group size on the model card, estimate the weights from the repository's file size, and add headroom for the cache and activations. The VRAM calculator will check a specific model and context length. Confirm that the serving software you plan to use loads that GPTQ layout for that architecture before you rent the card, because support differs between runtimes and between model families. If you want to compare against a 16-bit baseline on a larger card, Aquanode rents GPUs by the hour, so the comparison costs a short rental rather than a purchase.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
AWQ
AWQ (Activation-aware Weight Quantization) compresses LLM weights to low-bit integers by protecting the channels that activations show matter most.
GGUF
GGUF is a single-file binary format that packs a model's weights and metadata together, used by GGML-based runtimes such as llama.cpp.
EXL2
EXL2 is the quantization format of ExLlamaV2: GPTQ-based weights with mixed bit widths so a model can hit any average from 2 to 8 bits.
FP8
FP8 is an 8-bit floating-point format (E4M3 and E5M2) that halves memory versus 16-bit while keeping an exponent, used to run LLMs on newer GPUs.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.