How to fine-tune Llama 3.1 (LoRA, QLoRA)
Fine-tune Llama 3.1 with LoRA or QLoRA: which checkpoint to start from, the chat template, and what the Llama 3.1 license requires of derivatives.
See the Llama 3 hub for sizes. Parameter-efficient methods (LoRA and QLoRA) are the usual way to adapt these models on a single GPU. Estimate cost with the training cost calculator or read the fine-tuning overview.
Which checkpoint
- Start from the Instruct checkpoint to adjust style, format or domain while keeping chat behavior. Meta describes it as tuned with supervised fine-tuning and RLHF.
- Start from the base checkpoint for large-scale or heavily different instruction data, where you supply the whole post-training recipe.
The cards document no LoRA target modules, so choose them with your training library's defaults for Llama architectures.
Chat template
Train with the same template inference uses: format examples with the tokenizer's apply_chat_template. The Instruct card says tool use is supported through that template, so tool-calling data should be rendered the same way.
License limits on derivatives
From the Llama 3.1 Community License text:
- If you distribute a model trained or improved using Llama materials, include "Llama" at the beginning of its name.
- Prominently display "Built with Llama" on a related website, interface, blog post, about page or product documentation.
- Keep the notice "Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved" in a Notice file.
- Licensees with more than 700 million monthly active users must request a license from Meta.
- Use must follow Meta's Acceptable Use Policy.
Other notes
- Official languages are eight (English, German, French, Italian, Portuguese, Hindi, Spanish, Thai). Fine-tuning data in other languages needs its own evaluation.
Memory to fine-tune Llama 3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Meta-Llama-3-8B-Instruct | 8.0B | 120 GBRTX A5000 × 5 · $0.880/hr | 15.0 GBRTX A4000 · $0.167/hr | 3.9 GBRTX 4070 Super · $0.121/hr | 4.8 × 10^19 FLOPs |
| Meta-Llama-3-70B | 70.6B | 1051 GBNo live fit | 131 GBRTX A5000 × 6 · $1.06/hr | 33.9 GBRTX A6000 · $0.363/hr | 4.2 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct
- https://huggingface.co/meta-llama/Llama-3.1-8B
- https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct/blob/main/LICENSE
Updated 2026-10-07.