How to fine-tune Llama 3.1 (LoRA, QLoRA)
Base vs instruct checkpoints, chat template and Llama 3.1 Community License terms to check before fine-tuning Llama 3.1 with LoRA or QLoRA.
This guide covers adapting Llama 3.1. For adapter methods see LoRA and QLoRA. To size a run use the training cost calculator or the fine-tuning overview.
Base or instruct
The 8B Instruct card says the model is tuned with supervised learning and RLHF for dialogue, and that a separate pretrained base exists at meta-llama/Llama-3.1-8B. Start from Instruct to keep chat and tool behavior, or from the base model if you are training your own format.
Chat template
Train with the template shipped in the tokenizer. The card shows tools passed through tokenizer.apply_chat_template(), so tool-use examples should be rendered with the same call. It requires Transformers 4.43.0 or newer. LoRA target modules are not documented on the card.
License limits on derivatives
Under the Llama 3.1 Community License as described on the card:
- "Llama" must appear at the beginning of any derivative AI model name.
- "Built with Llama" must be shown prominently.
- A license from Meta is needed above 700 million monthly active users.
- The license agreement and copyright notice must be included when redistributing.
Read the full license before shipping a fine-tuned model.
Memory to fine-tune Llama 3.1, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | 8.0B | 120 GBRTX A5000 × 5 · $0.880/hr | 15.0 GBRTX A4000 · $0.167/hr | 3.9 GBRTX 4070 Super · $0.121/hr | 4.8 × 10^19 FLOPs |
| Llama-3.1-70B-Instruct | 70.6B | 1051 GBNo live fit | 131 GBRTX A5000 × 6 · $1.06/hr | 33.9 GBRTX A6000 · $0.363/hr | 4.2 × 10^20 FLOPs |
| Llama-3.1-405B | 405.9B | 6048 GBNo live fit | 756 GBRTX PRO 6000 × 8 · $11.00/hr | 195 GBRTX A6000 × 5 · $1.81/hr | 2.4 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.