LoRA Fine-Tuning Guide: Rank, QLoRA and VRAM (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

LoRA fine-tuning freezes a pre-trained model and trains two small matrices per targeted layer, so you update well under 1% of the weights and skip almost all optimizer memory. QLoRA goes one step further and stores the frozen base in 4-bit, which is why a 65B model can be fine-tuned on a single 48 GB GPU according to the QLoRA paper.

This guide covers the mechanism, the hyperparameters that matter (rank, alpha, target modules, learning rate), a runnable QLoRA setup, and a VRAM table computed from a formula you can reuse. It is part of our LLM fine-tuning frameworks guide.

TL;DR

  • LoRA learns a low-rank update B x A next to each frozen weight matrix. Only A and B get gradients and optimizer state.
  • The expensive part of full fine-tuning is optimizer state, not the weights. LoRA removes most of it, but the frozen base still has to fit in VRAM.
  • QLoRA stores that frozen base in 4-bit (NF4), cutting base memory to roughly a quarter of BF16. If the BF16 base does not fit, switch to QLoRA before buying a bigger GPU.
  • Start with rank 16, target_modules="all-linear", and a learning rate around ten times what you would use for full fine-tuning.
  • LoRA matches full fine-tuning on small and medium supervised datasets and reinforcement learning, and can fall behind when the dataset is large relative to adapter capacity.

If you are still deciding whether to tune at all, read what AI model fine-tuning involves first.

How LoRA works

The LoRA paper (Hu et al., 2021) freezes the pre-trained weights and injects trainable rank decomposition matrices into the transformer layers. For a frozen weight matrix W0 of shape d by k, LoRA learns B (d by r) and A (r by k) with a small rank r, and the layer uses the sum of the frozen path and the adapter path.

h = W0 x + (alpha / r) * B A x

B starts at zero, so at step 0 the adapter is a no-op and the model behaves exactly like the base. PEFT documents this as the default initialization (init_lora_weights=True, with B set to zero).

The parameter saving is easy to compute. For one square 4096 by 4096 matrix at rank 16:

full matrix:   4096 * 4096        = 16,777,216 params
LoRA adapter:  16 * (4096 + 4096) =    131,072 params   (0.78%)

That is computed arithmetic, not a benchmark. The paper's own headline numbers, measured on GPT-3 175B against Adam fine-tuning, are a 10,000 times reduction in trainable parameters and a 3 times reduction in GPU memory. It reports training VRAM falling from 1.2 TB to 350 GB, and, with r = 4 and only the query and value projections adapted, a checkpoint shrinking from 350 GB to 35 MB. It also reports a 25% training speedup on that model and no added inference latency, because the adapter can be merged into W0 after training. Those figures are the authors', on GPT-3.

Because the adapter is a separate small file, one frozen base can serve many tasks. The paper's arithmetic: storing 100 adapted models costs about 354 GB with LoRA versus about 35 TB as 100 full copies.

For the short definition, see the glossary entries for LoRA and QLoRA.

Where the memory goes

Full fine-tuning with mixed-precision Adam needs roughly 16 bytes per parameter for model states, per the ZeRO paper: 2 bytes for the half-precision weights, 2 for gradients, and 12 for the fp32 optimizer states (K = 12). That excludes activations.

LoRA keeps the frozen base at its storage precision (2 bytes per parameter in BF16) and pays the 16 bytes only on the tiny adapter. QLoRA stores the base at 0.5 bytes per parameter. Three formulas, all computed:

full fine-tune  = 16 * P bytes
LoRA (BF16 base) = 2 * P bytes + adapter states
QLoRA (4-bit)    = 0.5 * P bytes + adapter states + quantization constants

Adapter states are small. If you adapt four square 4096 by 4096 projections in each of 32 layers at rank 16 (an illustrative layout, not a specific model), that is 32 _ 4 _ 131,072 = 16,777,216 adapter parameters. At 16 bytes each (fp32 weights, gradients and Adam states) that is about 0.27 GB.

Base modelFull fine-tune (16 B/param)LoRA, BF16 base (2 B/param)QLoRA, 4-bit base (0.5 B/param)
8B128 GB16 GB4 GB
32B512 GB64 GB16 GB
70B1,120 GB140 GB35 GB

All values are computed from the formulas above and cover model states only. Activations, which grow with batch size and sequence length, come on top, and gradient checkpointing trades compute for less of them. The QLoRA paper's 65B on 48 GB claim is consistent with this math: 65 billion times 0.5 bytes is 32.5 GB, leaving room for adapters and activations.

Mapping that to cards (VRAM from NVIDIA's spec pages: RTX 4090 24 GB, L40S 48 GB, H100 SXM 80 GB, H200 141 GB):

  • 8B QLoRA or LoRA: fits a 24 GB card. LoRA at 16 GB for the base leaves about 8 GB for activations, so keep batch and sequence length modest.
  • 32B QLoRA: 16 GB of base fits 24 GB with a short context, comfortably on 48 GB.
  • 70B QLoRA: 35 GB base fits 48 GB and 80 GB cards.
  • 70B LoRA in BF16: 140 GB of base does not fit one 80 GB GPU. Use multiple GPUs (with FSDP sharding) or QLoRA.
  • Full fine-tuning at 8B: 128 GB of states already exceeds one 80 GB card, which is why full runs shard across GPUs.

You can check any card with our VRAM pages and the broader VRAM sizing guide.

QLoRA: the 4-bit base

QLoRA (Dettmers et al., 2023) backpropagates through a frozen, 4-bit quantized model into LoRA adapters. The paper introduces three pieces: NF4 (4-bit NormalFloat, a data type it calls information theoretically optimal for normally distributed weights), double quantization (quantizing the quantization constants to cut the average footprint), and paged optimizers (to absorb memory spikes). The authors report that their Guanaco family reached 99.3% of ChatGPT's performance level on the Vicuna benchmark after 24 hours of fine-tuning on a single GPU. That is their benchmark and their evaluation setup.

The weights are stored in 4-bit but computed in a higher precision. The Hugging Face PEFT docs show the standard configuration:

import torch
from transformers import BitsAndBytesConfig

config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)

For background on what 4-bit storage does to accuracy, see quantization.

A runnable setup

The PEFT quantization guide gives this sequence for a QLoRA model:

from transformers import AutoModelForCausalLM
from peft import prepare_model_for_kbit_training, LoraConfig, get_peft_model

model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.1", quantization_config=config)
model = prepare_model_for_kbit_training(model)

lora = LoraConfig(
    r=16,
    lora_alpha=8,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora)

If you would rather not wire the loop yourself, TRL's SFTTrainer accepts a peft_config directly (and a quantization_config for QLoRA). This is the pattern from the TRL docs:

from datasets import load_dataset
from trl import SFTTrainer
from peft import LoraConfig

dataset = load_dataset("trl-lib/Capybara", split="train")

trainer = SFTTrainer(
    "Qwen/Qwen3-0.6B",
    train_dataset=dataset,
    peft_config=LoraConfig(),
)
trainer.train()

TRL's docs note that when training adapters you typically use a higher learning rate, around 1e-4, than the SFTConfig default of 2e-5. For more on the trainer see our TRL library guide. If you want a faster wrapper around the same idea, see the Unsloth guide, and for a config-file workflow see LLaMA-Factory.

Choosing the hyperparameters

PEFT's LoraConfig defaults are r=8, lora_alpha=8, lora_dropout=0.0, bias="none", and no target modules unless the architecture is known.

Rank (r). Rank sets adapter capacity. The paper reports that on GPT-3, small ranks already suffice for many tasks, and that the update has low intrinsic rank. Raise r when the dataset is large or the task is far from the base model's strengths. A later study, "LoRA Learns Less and Forgets Less," found that in standard low-rank settings LoRA substantially underperforms full fine-tuning on programming and mathematics, and that full fine-tuning learns perturbations with a rank 10 to 100 times greater than typical LoRA configurations. So treat rank as a knob to test, not a constant.

Alpha. Output is scaled by alpha divided by r. A common convention is to keep alpha and r in a fixed ratio so changing rank does not silently change the effective step size. With use_rslora=True, PEFT instead scales by alpha divided by the square root of r, following the rank-stabilized LoRA paper, whose claim is that the standard factor slows learning at higher ranks.

Target modules. PEFT's default for known architectures is the query and value projections. The PEFT docs note that QLoRA, which adapts all linear layers, can provide performance equal to a fully fine-tuned model, and recommend target_modules="all-linear". Thinking Machines' "LoRA Without Regret" study makes the same case from experiments: apply LoRA to all layers, especially the MLP layers, and attention-only adapters give no extra benefit over MLP-only.

Learning rate. The same study found the optimal learning rate for LoRA is consistently about 10 times the full fine-tuning rate, and TRL's docs echo the higher-rate advice for adapters.

Dropout. Small values such as 0.05 are common in the PEFT examples. The default is zero.

LoRA versus full fine-tuning

The honest answer from the literature is: it depends on data size and task.

  • "LoRA Without Regret" reports that for small to medium supervised datasets, LoRA learns with the same sample efficiency as full fine-tuning and reaches the same final performance when the details above are right. When the dataset exceeds adapter capacity, LoRA underperforms. At large batch sizes it pays a larger loss penalty, and raising rank does not fix that.
  • The same study reports that LoRA matches full fine-tuning in policy-gradient reinforcement learning even at rank 1, which they attribute to RL absorbing very little information per episode. That is relevant to GRPO.
  • "LoRA Learns Less and Forgets Less" found a gap on code and math, but also less forgetting of the base model's abilities.

Practical rule: begin with LoRA or QLoRA. Move to full fine-tuning only when you have a large dataset, a measured quality gap, and the memory (see the table) to shard it.

Merging and serving the adapter

After training you have two options. Keep the adapter separate and load it onto the base, which lets one base serve many adapters, or fold it in. The PEFT LoRA guide shows both:

from transformers import AutoModelForCausalLM
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.1")
model = PeftModel.from_pretrained(base_model, "alignment-handbook/zephyr-7b-sft-lora")
model = model.merge_and_unload()

merge_and_unload() is not in-place, so assign the result. Use merge_adapter() instead if you want to unmerge later. For serving, see serving LLMs with vLLM.

Run it on a cloud GPU

A 24 GB card handles 8B QLoRA, a 48 GB card handles 70B QLoRA, and an 80 GB card gives you headroom for larger batches. The live box shows what is available right now.

FAQ

What rank should I use for LoRA?

PEFT defaults to 8, and 16 is a common starting point. Treat it as a hyperparameter: raise it if the adapter underfits a large dataset, and compare against a lower rank on a held-out set before paying for more.

How much VRAM does QLoRA need?

Computed as 0.5 bytes per base parameter plus adapter states and activations: about 4 GB for the 8B base, 16 GB for 32B, and 35 GB for 70B before activations. The QLoRA paper reports a 65B model on a single 48 GB GPU.

Is LoRA as good as full fine-tuning?

On small and medium supervised datasets and on RL, the "LoRA Without Regret" study reports parity. On large code or math datasets with low ranks, "LoRA Learns Less and Forgets Less" reports full fine-tuning ahead.

Can I fine-tune a 70B model on one GPU?

With QLoRA, yes on a 48 GB or 80 GB card for the base weights (35 GB computed), with short sequences. In BF16 LoRA the 140 GB base needs multiple GPUs.

Does LoRA slow down inference?

Not once merged. The paper states LoRA adds no extra inference latency because the adapter can be merged into the base weights. Unmerged adapters add a small extra matrix multiply.

Should I use DoRA instead?

DoRA is a LoRA variant that adds a learned magnitude. PEFT notes it has more overhead than plain LoRA. See our PEFT methods guide.

Sources

#fine-tuning#llm training#lora#qlora#peft#vram

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.