What is LoRA?

Abbreviated LoRA

LoRA (Low-Rank Adaptation) is a fine-tuning method that freezes a pre-trained model's weights and trains a small number of extra weights injected into each layer. It was introduced in "LoRA: Low-Rank Adaptation of Large Language Models" (Hu et al., submitted June 2021). The paper describes it as freezing the pre-trained weights and injecting trainable rank decomposition matrices into the transformer layers.

What the paper claims

Compared with fine-tuning GPT-3 175B with Adam, the abstract reports that LoRA cuts the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times. It says LoRA performs on par with or better than full fine-tuning on RoBERTa, DeBERTa, GPT-2 and GPT-3, with higher training throughput, and that unlike adapter layers it adds no extra inference latency. These numbers are the authors' own, measured on those models and tasks.

How it works

A weight matrix W in a layer is left untouched. LoRA learns two much smaller matrices whose product has the same shape as W, and adds that product to the layer's output. The size of the smaller matrices is the rank, which you choose: a lower rank means fewer trainable parameters. Only the small matrices receive gradients and optimizer state, so most of the memory that full fine-tuning spends on optimizer state disappears.

Because the learned update is a separate small file (the adapter), one frozen base model can serve many tasks. Each task can get its own adapter (the abstract states LoRA adds no extra inference latency).

LoRA and the base model's precision

LoRA does not shrink the frozen base. The full base model still has to sit in GPU memory during training, in BF16 or whatever precision you load it in, plus activations. That is the gap QLoRA fills: it keeps the same adapters but stores the frozen base in 4-bit, using quantization. If the BF16 base does not fit, move to QLoRA before reaching for a larger GPU.

LoRA in the catalog

The catalog currently has no model whose name carries "LoRA", because adapters are usually published as small separate files rather than full model pages. What you will see are full models that were produced by fine-tuning, and their cards say how. The method applies to any of the transformer models listed on Aquanode, and the base-model pages tell you the starting size to plan around. See what AI model fine-tuning involves for when to tune at all.

What it means when you pick a GPU

Size the job as base weights plus adapter and optimizer state plus activations. The adapter and its optimizer state are small, so the base model and your sequence length dominate VRAM. Activations grow with batch size and sequence length and the sequence length. For models too large for one card, FSDP can shard the frozen base across several. Use the VRAM calculator to check your model and sequence length. Aquanode rents GPUs by the hour, which fits a fine-tune that runs for a few hours.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.