What is GPTQ?

Abbreviated GPTQ

GPTQ is a post-training quantization method that compresses the weights of a generative pre-trained transformer to a few bits each, in one shot, without retraining. It comes from GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers by Elias Frantar, Saleh Ashkboos, Torsten Hoefler and Dan Alistarh, first submitted in October 2022. The paper describes it as "a new one-shot weight quantization method based on approximate second-order information."

What the paper claims

The authors report three results, all from the paper's own experiments:

  • A model with 175 billion parameters can be compressed to 3 to 4 bits per weight while keeping accuracy.
  • The quantization process takes approximately four GPU hours for that model.
  • The compressed 175-billion-parameter model fits inside a single GPU for generative inference, and inference runs about 3.25x faster on high-end GPUs and 4.5x faster on more affordable ones, compared with the baseline the paper uses.

"One-shot" means the weights are rounded in a single pass over a small amount of data, not by training. "Second-order information" means the method uses curvature of the error, rather than just rounding each weight to its nearest value, to decide how to adjust the remaining weights as each one is rounded. See quantization for the general trade between precision and memory.

GPTQ in practice

Like AWQ, GPTQ names an algorithm, not a single file format. Hugging Face repositories labelled GPTQ carry weights produced by it, usually with the bit width in the name (for example Int4). Related names appear too: the EXL2 format is described by its own README as based on the same optimization method as GPTQ, and some repositories use a GPTQ-compatible layout produced by other tools. The GGUF format is a separate container with its own quantization types, so a GGUF file is not a GPTQ file.

Example models on Aquanode

  • Vishva007 Gemma 4 E4B it W4A16 AutoRound GPTQ is a 4-bit build of Google's Gemma 4 E4B instruction-tuned model. Its card describes INT4 weights with BF16 activations (W4A16), a group size of 128, quantization with AutoRound, and calibration on 256 samples at 2048 sequence length. Only the language-model layers are quantized to INT4; the vision tower, audio tower and projectors stay in BF16. The model is subject to the Gemma Terms of Use.

This card is a useful illustration of two points. A repository named GPTQ can be produced by a different tool than the original GPTQ code, so read the card for what was actually done. And quantizing only part of a multimodal model leaves the rest at full size, which affects the memory you need.

What GPTQ does not shrink

GPTQ lowers the weight storage. It does not shrink the KV cache, which grows with context length and the number of requests, and it does not change the activations a runtime computes in higher precision. Weights stored in 4-bit are converted when the math runs, so the savings are in memory and bandwidth rather than in the arithmetic itself.

What it means when you pick a GPU

Read the bit width and group size on the model card, estimate the weights from the repository's file size, and add headroom for the cache and activations. The VRAM calculator will check a specific model and context length. Confirm that the serving software you plan to use loads that GPTQ layout for that architecture before you rent the card, because support differs between runtimes and between model families. If you want to compare against a 16-bit baseline on a larger card, Aquanode rents GPUs by the hour, so the comparison costs a short rental rather than a purchase.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.