What is AWQ?
Abbreviated AWQ
AWQ, short for Activation-aware Weight Quantization, is a method for compressing a large language model by storing its weights in low-bit integers. It was introduced in a paper by Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan and Song Han, first submitted in June 2023 as AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. The paper received the MLSys 2024 Best Paper Award.
The core idea
The paper's key insight is that "not all weights in an LLM are equally important." Protecting only 1% of salient weights can greatly reduce quantization error, the authors write. The twist is how those weights are found. Rather than looking at the weights themselves, AWQ looks at the activation distributions to decide which channels matter most.
The paper then derives mathematically that scaling up the important channels reduces quantization error, and uses that scaling as a transformation applied before the weights are rounded. It does this without backpropagation, so quantizing a model does not involve training it.
This is a form of post-training quantization: you start from a finished model and produce a smaller one, rather than training at low precision. The paper reports results across language modeling, coding, math benchmarks and multimodal models.
Speed as well as size
The authors also released TinyChat, an inference framework for running AWQ models. They report more than 3x speedup over the Hugging Face FP16 implementation on both desktop and mobile GPUs, and that it lets larger models run on mobile hardware. Those are the paper's own measurements on its own setup. Real speed depends on your serving software, the batch size and the card, so treat it as a direction rather than a promise.
How AWQ repositories look
AWQ is an algorithm, not a file format, so AWQ models are published as ordinary Hugging Face repositories in whatever serving stack's layout the quantizer chose. You will see "AWQ" in the name, often with a bit width such as INT4. In the Aquanode catalog, many of these are re-quantized copies of an existing model made by a third party. Compare that to GGUF, where the file format itself is the thing named, and to GPTQ, the other widely used post-training method for GPUs.
Example models on Aquanode
- cyankiwi Gemma 4 E4B it AWQ INT4 is a 4-bit AWQ build of Google's Gemma 4 E4B instruction-tuned model. Its card lists text, image and audio input and a 128K-token context window.
- QuantTrio Qwen3.5 9B AWQ is an AWQ build of Qwen3.5-9B under the Apache 2.0 license. Its card shows serving it with
vllm serve QuantTrio/Qwen3.5-9B-AWQ, and lists SGLang and Hugging Face Transformers as other options. - QuantTrio Qwen3.5 4B AWQ is the same treatment for the 4B model, also Apache 2.0, whose card describes a native context length of 262,144 tokens.
- cyankiwi Ornith 1.0 35B AWQ FP8 is a quantized build of Ornith-1.0-35B, a 35-billion-parameter model for agentic coding built on the Qwen 3.5 architecture and licensed MIT, per its card. The name combines AWQ with FP8, so read the card for which part of the model uses which.
What AWQ does not shrink
AWQ reduces the weights. The KV cache is a separate memory cost that grows with context length and batch size, and quantizing weights does not change it. A 4-bit model with a very long context can still need a lot of memory.
What it means when you pick a GPU
An AWQ repository lets a model that is too large for a card in 16-bit fit in far less VRAM, and the model card tells you the bit width. Add room for the cache and activations, then check the exact model with the VRAM calculator. Make sure the runtime you plan to use supports AWQ for that architecture, since a new model family can arrive before every server handles its quantized form. Aquanode rents GPUs by the hour, which makes it practical to try a quantized build on a smaller card first.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
GPTQ
GPTQ is a one-shot post-training method that compresses LLM weights to 3 or 4 bits using approximate second-order information.
GGUF
GGUF is a single-file binary format that packs a model's weights and metadata together, used by GGML-based runtimes such as llama.cpp.
EXL2
EXL2 is the quantization format of ExLlamaV2: GPTQ-based weights with mixed bit widths so a model can hit any average from 2 to 8 bits.
FP8
FP8 is an 8-bit floating-point format (E4M3 and E5M2) that halves memory versus 16-bit while keeping an exponent, used to run LLMs on newer GPUs.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.