How to fine-tune Llama 2 (LoRA, QLoRA)
Chat format and Llama 2 Community License limits (attribution, 700M MAU, no improving other LLMs) before fine-tuning.
This guide covers adapting the Llama 2 family. Memory per method and GPU counts appear below, and the training cost calculator estimates runtime. For method background see fine-tuning, LoRA and QLoRA.
Setup
- Access is gated: accept Meta's license on Hugging Face and authenticate before downloading.
- If you train from the chat checkpoint, format examples with its prompt layout:
INSTand<<SYS>>tags with BOS and EOS tokens, andstrip()your inputs to avoid double spaces, as the card advises. - Base and chat checkpoints exist as separate repos. Pick chat to keep dialogue behavior.
- The card lists a 4,096 token context, so keep training sequences within it.
- The card does not document LoRA target modules, so use your trainer's defaults.
- The card says testing was English only. Expect to evaluate other languages yourself.
License
The weights are under the LLAMA 2 Community License. Terms that affect derivatives:
- Distributing the Llama Materials requires including the attribution notice: "Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved."
- If your products or services had more than 700 million monthly active users in the month before the Llama 2 release date, you must request a license from Meta.
- You may not use the Llama Materials or their output to improve any other large language model (excluding Llama 2 or derivative works thereof).
- The Acceptable Use Policy applies to the model and to derivatives, and includes a prohibition on military and warfare applications.
Read the license text to confirm how it treats your fine-tuned adapter or merged model.
Memory to fine-tune Llama 2, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Llama-2-7b-hf | 6.7B | 100 GBRTX A5000 × 5 · $0.880/hr | 12.6 GBRTX A4000 · $0.167/hr | 3.2 GBRTX 4070 Super · $0.121/hr | 4.0 × 10^19 FLOPs |
| Llama-2-13b-chat-hf | 13.0B | 194 GBRTX A6000 × 5 · $1.81/hr | 24.2 GBRTX 4080 Super · $0.338/hr | 6.3 GBRTX 4070 Super · $0.121/hr | 7.8 × 10^19 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/meta-llama/Llama-2-7b-chat-hf
- https://huggingface.co/meta-llama/Llama-2-7b-chat-hf/blob/main/LICENSE.txt
Updated 2026-10-07.