How to fine-tune Llama 3.1 (LoRA, QLoRA)

Fine-tune Llama 3.1 with LoRA or QLoRA: which checkpoint to start from, the chat template, and what the Llama 3.1 license requires of derivatives.

See the Llama 3 hub for sizes. Parameter-efficient methods (LoRA and QLoRA) are the usual way to adapt these models on a single GPU. Estimate cost with the training cost calculator or read the fine-tuning overview.

Which checkpoint

  • Start from the Instruct checkpoint to adjust style, format or domain while keeping chat behavior. Meta describes it as tuned with supervised fine-tuning and RLHF.
  • Start from the base checkpoint for large-scale or heavily different instruction data, where you supply the whole post-training recipe.

The cards document no LoRA target modules, so choose them with your training library's defaults for Llama architectures.

Chat template

Train with the same template inference uses: format examples with the tokenizer's apply_chat_template. The Instruct card says tool use is supported through that template, so tool-calling data should be rendered the same way.

License limits on derivatives

From the Llama 3.1 Community License text:

  • If you distribute a model trained or improved using Llama materials, include "Llama" at the beginning of its name.
  • Prominently display "Built with Llama" on a related website, interface, blog post, about page or product documentation.
  • Keep the notice "Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved" in a Notice file.
  • Licensees with more than 700 million monthly active users must request a license from Meta.
  • Use must follow Meta's Acceptable Use Policy.

Other notes

  • Official languages are eight (English, German, French, Italian, Portuguese, Hindi, Spanish, Thai). Fine-tuning data in other languages needs its own evaluation.

Memory to fine-tune Llama 3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Meta-Llama-3-8B-Instruct8.0B120 GBRTX A5000 × 5 · $0.880/hr15.0 GBRTX A4000 · $0.167/hr3.9 GBRTX 4070 Super · $0.121/hr4.8 × 10^19 FLOPs
Meta-Llama-3-70B70.6B1051 GBNo live fit131 GBRTX A5000 × 6 · $1.06/hr33.9 GBRTX A6000 · $0.363/hr4.2 × 10^20 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.