What is EXL2?

Abbreviated EXL2

EXL2 is the quantization format used by ExLlamaV2, an inference library for running language models on consumer GPUs. According to the ExLlamaV2 README, "EXL2 is based on the same optimization method as GPTQ and supports 2, 3, 4, 5, 6 and 8-bit quantization." Its distinguishing feature is that one model can mix several of those bit widths.

Mixed bit widths

The README explains that the format allows mixing quantization levels within a model to achieve any average bitrate between 2 and 8 bits per weight. Different levels can be applied within the linear layers, so parts of the model that matter more can be stored with more bits and others with fewer. The result is a file whose average is whatever you ask for, such as 4.65 bits per weight, instead of being limited to a fixed menu like 4-bit or 8-bit.

That flexibility is the practical point. With a fixed 4-bit format, a model either fits your card or it does not. With a continuous average, you can choose the highest quality that fits the memory you have. The README gives the example of running a Llama 2 70B model on a single 24 GB GPU.

How it relates to other formats

  • GPTQ is the underlying optimization method. EXL2 builds on it and adds the mixed-bitrate capability, so it is not simply a GPTQ file.
  • AWQ is a different algorithm, which protects salient channels found from activation statistics. The two are not interchangeable files.
  • GGUF is a single-file container for GGML-based runtimes with its own quantization types. EXL2 is read by ExLlamaV2, not by those runtimes.

For the general idea of trading precision for memory, see quantization.

Project status

The ExLlamaV2 repository carries a notice that "This project is archived for now. Development continues on ExLlamaV3." It is released under the MIT license. An archived project still works, but new model architectures are not guaranteed to arrive for it, so check that the model you want is supported before you plan around EXL2. EXL2 repositories on Hugging Face are published by third parties who quantize existing models, and the repository name usually states the average bits per weight.

Example models on Aquanode

The Aquanode catalog does not currently list a model whose name carries the EXL2 label, so there is no model page to link here. If you need a model in EXL2 form, quantize one yourself with the ExLlamaV2 tools on a rented GPU, or pick an AWQ or GPTQ build, which the catalog does list.

What EXL2 does not shrink

Like any weight-only scheme, EXL2 lowers weight storage and leaves the KV cache alone. The cache grows with context length and with how many requests you run at once, so a model that fits on paper at 4 bits can still run out of memory at a long context.

What it means when you pick a GPU

The average bits per weight in the repository name gives you a direct handle: multiply the parameter count by that average and divide by eight to estimate weight size, then add cache and activations. The README's 24 GB example shows the intended use, squeezing a large model onto one consumer-class card. Check the VRAM calculator for your model and context length, and confirm the library supports the architecture. Aquanode rents GPUs by the hour, so you can test a given bitrate on a card before settling on it.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.