What is QLoRA?
Abbreviated QLoRA
QLoRA is a fine-tuning method that freezes a language model's weights in 4-bit precision and trains only a small set of extra weights, called LoRA adapters, on top of them. The base model is quantized, so it uses roughly a quarter of the memory it would at 16-bit, and only the adapters get gradients and optimizer state. The paper that introduced it (Dettmers et al., 2023) reports fine-tuning a 65-billion-parameter model on a single 48GB GPU, where full 16-bit fine-tuning of the same model needs more than 780 GB.
The three pieces
- 4-bit NormalFloat (NF4). A 4-bit data type designed for weights that are roughly normally distributed. The base model is stored in NF4 and converted to BF16 on the fly for each matrix multiplication, so quantization saves memory while the math still runs in 16-bit.
- Double quantization. Quantization stores a scale constant for every block of 64 weights. QLoRA quantizes those constants too. On average this takes the overhead from 0.5 bits per parameter to 0.127 bits, saving about 0.37 bits per parameter (roughly 3 GB for a 65B model, per the paper).
- Paged optimizers. They use NVIDIA's unified memory feature to move optimizer state to CPU memory automatically when the GPU runs out and back when it is needed, which absorbs the memory spikes that long sequences cause during gradient checkpointing.
LoRA itself is the other half. Instead of updating a weight matrix W, it learns two small matrices whose product is added to W's output, and trains only those. The paper applies adapters to every linear layer of the transformer, because it found that LoRA on all linear layers is needed to match 16-bit full fine-tuning.
Worked example: Llama 3.1 8B
Llama 3.1 8B has 8.03 billion parameters (computed from its published config.json: 32 layers, hidden size 4096, feed-forward size 14,336, 8 key/value heads, 128,256-token vocabulary).
- Frozen base weights. With NF4 plus double quantization at 4.127 bits per parameter (4 bits plus 0.127), that is 8.03 billion x 4.127 / 8 = 4.14 GB. In BF16 the same weights take 16.06 GB.
- Adapters. At LoRA rank 16 on all seven linear layers in each transformer block (the paper's own settings use rank 64, so this is smaller), each block has 1,310,720 trainable parameters, or 41.9 million across 32 layers. In fp32 with gradients and Adam moments, that is 16 bytes each: 0.67 GB.
- Full fine-tuning for comparison. The 16 bytes per parameter that the ZeRO paper counts for mixed-precision Adam gives 8.03 billion x 16 = 128.5 GB of model state before activations.
So the persistent memory is about 4.8 GB with QLoRA against 128.5 GB for full fine-tuning, which is 27 times less. What is left is activations. The paper's own breakdown for a 7B model at batch size 1 is a useful reminder: the LoRA parameters took 26 MB, the input gradients 567 MB without gradient checkpointing and 18 MB per sequence with it, and the 4-bit base model 5,048 MB. Activations and sequence length decide what is left, not the adapters, so gradient checkpointing matters.
What it costs
QLoRA trades memory for extra work: every matrix multiply includes a dequantization step that LoRA on a 16-bit base model does not have. The paper reports that NF4 with double quantization matched 16-bit LoRA performance in its tests, but it measured specific models and tasks, so check on yours. The adapters were trained against the quantized base, so evaluate the result in the form you plan to serve it.
What it means when you pick a GPU
Size the frozen base at about half a byte per parameter plus 0.127 bits, then add adapter state and activations. On that basis:
- A 7B to 8B model needs roughly 5 GB for weights and adapters, leaving room for activations on a 16GB card. An RTX 4090 with 24GB leaves comfortable headroom for longer sequences.
- A 70B model is 70.55 billion x 4.127 / 8 = 36.4 GB of base weights, which fits a 48GB card (such as an L40S) with about 11 GB for adapters, activations and the CUDA context. That is tight, so use gradient checkpointing and moderate sequence lengths. An 80GB H100 gives far more room.
- QLoRA runs on one GPU, which avoids interconnect questions altogether. When you do need several GPUs, FSDP can shard a QLoRA model, since PyTorch's FSDP2 documentation lists NF4 support.
Use the VRAM calculator to check your exact model and sequence length, and see what AI model fine-tuning involves for when to tune at all. Aquanode rents GPUs by the hour, which suits a fine-tune that runs for a few hours.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
FSDP (Fully Sharded Data Parallel)
FSDP shards a model's weights, gradients and optimizer state across GPUs and gathers each layer only when needed, so models too big for one GPU can train.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.
Unified Memory
Unified memory lets the CPU and GPU share one memory pool, as in Apple M-series chips and NVIDIA Grace Hopper. It adds capacity, not bandwidth.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.