PEFT (parameter-efficient fine-tuning) means adapting a large pre-trained model by training only a small number of extra or selected parameters while the rest stays frozen. In practice, LoRA is the default, DoRA is its drop-in upgrade when you can afford a little overhead, and prompt-style methods are lighter but usually weaker on small models.
This post maps the main PEFT families, says what each one actually trains, shows the Hugging Face PEFT code, and tells you which to pick. It sits under our LLM fine-tuning frameworks guide.
TL;DR
- LoRA adds trainable low-rank matrices beside frozen weights and can be merged back at no inference cost. It is the default for a reason.
- DoRA splits each weight into magnitude and direction and applies LoRA to the direction. The authors report it consistently beats LoRA; the PEFT docs warn of extra overhead.
- Adapters, prefix tuning, prompt tuning, P-tuning and (IA)3 are older or lighter designs. They still matter for tiny parameter budgets and multi-task serving.
- AdaLoRA and VeRA shrink the budget further, by reallocating rank or by sharing one random pair of matrices across layers.
- Pick by constraint: memory (QLoRA), quality (LoRA on all linear layers, then DoRA), or parameter count (VeRA, prompt tuning).
What PEFT is, and what the library is
"PEFT" names two things. The idea: per the Hugging Face docs, PEFT methods fine-tune only a small number of (extra) model parameters, significantly decreasing computational and storage costs, while yielding performance comparable to a fully fine-tuned model. The library: the Hugging Face peft package, integrated with Transformers, Diffusers and Accelerate, which the docs describe as a framework for arbitrary adaptation methods (modifying weights, wrapping layers, manipulating KV-caches) and a reference implementation for many of them.
The library's PeftType enum lists the methods it supports. As of the current docs it includes PROMPT_TUNING, MULTITASK_PROMPT_TUNING, P_TUNING, PREFIX_TUNING, LORA, ADALORA, BOFT, IA3, LOHA, LOKR, OFT, XLORA, VERA, FOURIERFT, HRA, RANDLORA and more than a dozen newer entries. That list grows every release, which is why the practical question is not "which of 30 methods" but "which family".
For the end-to-end picture of why you fine-tune at all, read what AI model fine-tuning involves.
The families
PEFT methods fall into three broad families by where the new parameters live.
| Family | Where trainable parameters live | Examples | Can merge into base? |
|---|---|---|---|
| Reparameterization | Low-rank update to existing weights | LoRA, DoRA, AdaLoRA, VeRA | Yes |
| Addition (layers) | New small modules inserted into the network | Adapters | No, adds layers |
| Addition (inputs) | Learned vectors at the input or attention | Prompt tuning, prefix tuning, P-tuning | No |
| Scaling | Learned vectors that rescale activations | (IA)3 | Not covered here |
LoRA
LoRA freezes the pre-trained weights and trains a pair of low-rank matrices per targeted layer. The paper reports, for GPT-3 175B against Adam fine-tuning, 10,000 times fewer trainable parameters and 3 times less GPU memory, with no added inference latency. The full mechanism, memory math and rank guidance are in our LoRA fine-tuning guide. The one-line PEFT usage, from the docs:
from peft import LoraConfig, get_peft_model
config = LoraConfig(
r=16,
lora_alpha=16,
target_modules=["query", "value"],
lora_dropout=0.1,
bias="none",
modules_to_save=["classifier"],
)
model = get_peft_model(model, config)
model.print_trainable_parameters()
print_trainable_parameters() is the quickest check that you really are training a small fraction of the model. Defaults in LoraConfig are r=8, lora_alpha=8, lora_dropout=0.0.
If the frozen base is too large for your GPU, the memory-saving variant is QLoRA: same adapters over a 4-bit base. The QLoRA paper reports a 65B model on a single 48 GB GPU.
DoRA
DoRA (Weight-Decomposed Low-Rank Adaptation) decomposes each pre-trained weight into magnitude and direction and fine-tunes both, using LoRA only for the directional update so the trainable count stays small. The abstract says it consistently outperforms LoRA on fine-tuning LLaMA, LLaVA and VL-BART across downstream tasks, and that it improves both learning capacity and training stability without adding inference overhead. The abstract gives no single headline number, so treat the gain as task dependent and measure it on your data.
In PEFT it is one flag:
config = LoraConfig(use_dora=True, ...)
The PEFT docs list the trade-offs. DoRA introduces more overhead than plain LoRA during training, so they recommend merging the weights for inference. It supports only linear and Conv2D layers, and it does not work with mixed-adapter batches (adapter_names), where you must set use_dora=False. Also note that with quantized bases, PEFT's torchao notes say DoRA only works with int8 weight-only there.
When to use it: you already run LoRA on all linear layers, the rank is reasonable, and you still see a quality gap. Try the flag before raising rank.
LoRA variants worth knowing
PEFT exposes several drop-in LoRA variants through LoraConfig:
- rsLoRA (
use_rslora=True): scales adapters by alpha divided by the square root of the rank instead of alpha divided by rank. The paper argues the standard factor slows learning at higher ranks, so rsLoRA lets you use larger ranks productively. - PiSSA, OLoRA, LoftQ (
init_lora_weights): smarter adapter initialization. LoftQ initializes the adapter to minimize the quantization error of a quantized base, and the docs recommendtarget_modules="all-linear"with NF4 when you use it. - LoRA+: different learning rates for the A and B matrices, which the docs describe as reported to give up to 2 times faster fine-tuning and 1 to 2% better performance. That is the method authors' claim as quoted in PEFT.
AdaLoRA and VeRA
AdaLoRA observes that uniform ranks spend the budget evenly even though some weight matrices matter more. It parameterizes the update through singular value decomposition, allocates budget by importance and prunes singular values of less useful updates, and the abstract reports better results than baselines especially when the budget is low.
VeRA (Vector-based Random Matrix Adaptation) shares a single pair of low-rank matrices across all layers and learns small scaling vectors per layer. The abstract says it significantly reduces trainable parameters compared to LoRA at comparable performance, without a specific figure.
Both are attractive when you store hundreds of adapters, since adapter size dominates storage. For one or two tasks, the saving rarely matters.
Adapters, prefix tuning, prompt tuning, P-tuning, (IA)3
These are the older additive designs.
- Adapters (Houlsby et al., 2019) insert small trainable modules into a frozen network. For BERT on 26 text classification tasks, the paper reports coming within 0.4% of full fine-tuning on GLUE while adding about 3.6% parameters per task. Unlike LoRA, the inserted layers stay in the forward pass.
- Prefix tuning (Li and Liang) keeps the language model frozen and optimizes a small continuous task-specific prefix. The abstract reports training only 0.1% of parameters on GPT-2 and BART with performance comparable to fine-tuning at full data and better in low-data settings.
- Prompt tuning (Lester et al.) learns soft prompts at the input. The paper's key result is scale: once models pass a few billion parameters, soft prompts close the gap with full model tuning.
- P-tuning adds trainable continuous prompt embeddings next to discrete prompts, and the paper targets natural language understanding, reporting reduced instability across manual prompt choices.
- (IA)3 rescales activations with learned vectors, adding a tiny amount of new parameters. It underpins the T-Few recipe, which the paper says beat the prior state of the art on the RAFT benchmark by 6% absolute.
Practical read: prompt-style methods are strongest on very large models and weakest on small ones (the Lester result is about scale). For the 7B to 70B models most teams tune today, LoRA family methods are the common choice.
How to choose
| Constraint | Pick | Why |
|---|---|---|
| Base model barely fits in VRAM | QLoRA | 4-bit base, same adapters |
| Best quality for a single task | LoRA on all linear layers, then try DoRA | PEFT docs recommend all-linear; DoRA reported to beat LoRA |
| Many tasks, tiny adapter files | VeRA, AdaLoRA | Fewer trainable parameters per task |
| Very large frozen model, minimal training | Prompt tuning, prefix tuning, (IA)3 | Fraction of a percent of parameters |
| Aligning with preferences or rewards | LoRA with DPO or GRPO | TRL trainers take a peft_config |
The training loop is the same regardless of method. TRL's SFTTrainer and DPOTrainer both accept peft_config=LoraConfig(). Framework wrappers that apply these methods for you are covered in the Unsloth guide and the TRL library guide.
Run it on a cloud GPU
Most PEFT runs fit on a single 24 GB or 80 GB card. The live box shows what is rentable right now.
FAQ
What does PEFT stand for?
Parameter-efficient fine-tuning. It also names the Hugging Face peft library that implements the methods.
Is LoRA a PEFT method?
Yes, it is the most widely used one. It trains low-rank matrices next to frozen weights.
Is DoRA better than LoRA?
The DoRA authors report that it consistently outperforms LoRA on LLaMA, LLaVA and VL-BART tasks. PEFT's docs note more training overhead and fewer supported layer types, so benchmark it on your own data.
Which PEFT method uses the least memory?
Memory depends mostly on the frozen base, not the adapter. The biggest saving comes from quantizing the base (QLoRA). Among adapters, prompt tuning, prefix tuning and (IA)3 train the fewest parameters.
Can I combine PEFT with quantization?
Yes. The PEFT docs describe QLoRA as quantizing a model to 4 bits and training it with LoRA, and list VeRA, AdaLoRA and (IA)3 as also supporting bitsandbytes quantization.
Do PEFT adapters slow inference?
LoRA-style adapters can be merged into the base and add no latency. Adapters that insert layers, such as the original adapter design, stay in the forward pass.
Sources
- Hugging Face PEFT overview: https://huggingface.co/docs/peft/index
- PEFT types (supported methods list): https://huggingface.co/docs/peft/main/en/package_reference/peft_types
- PEFT LoRA reference (LoraConfig, DoRA, rsLoRA): https://huggingface.co/docs/peft/package_reference/lora
- PEFT LoRA developer guide (variants): https://huggingface.co/docs/peft/main/en/developer_guides/lora
- PEFT quantization guide: https://huggingface.co/docs/peft/main/en/developer_guides/quantization
- LoRA (Hu et al.): https://arxiv.org/abs/2106.09685
- QLoRA (Dettmers et al.): https://arxiv.org/abs/2305.14314
- DoRA (Liu et al.): https://arxiv.org/abs/2402.09353
- rsLoRA (Kalajdzievski): https://arxiv.org/abs/2312.03732
- AdaLoRA: https://arxiv.org/abs/2303.10512
- VeRA: https://arxiv.org/abs/2310.11454
- Adapters, Parameter-Efficient Transfer Learning for NLP (Houlsby et al.): https://arxiv.org/abs/1902.00751
- Prefix-Tuning (Li and Liang): https://arxiv.org/abs/2101.00190
- Prompt Tuning (Lester et al.): https://arxiv.org/abs/2104.08691
- P-Tuning (Liu et al.): https://arxiv.org/abs/2103.10385
- (IA)3 / T-Few (Liu et al.): https://arxiv.org/abs/2205.05638