How to fine-tune Llama 2 (LoRA, QLoRA)

Chat format and Llama 2 Community License limits (attribution, 700M MAU, no improving other LLMs) before fine-tuning.

This guide covers adapting the Llama 2 family. Memory per method and GPU counts appear below, and the training cost calculator estimates runtime. For method background see fine-tuning, LoRA and QLoRA.

Setup

  • Access is gated: accept Meta's license on Hugging Face and authenticate before downloading.
  • If you train from the chat checkpoint, format examples with its prompt layout: INST and <<SYS>> tags with BOS and EOS tokens, and strip() your inputs to avoid double spaces, as the card advises.
  • Base and chat checkpoints exist as separate repos. Pick chat to keep dialogue behavior.
  • The card lists a 4,096 token context, so keep training sequences within it.
  • The card does not document LoRA target modules, so use your trainer's defaults.
  • The card says testing was English only. Expect to evaluate other languages yourself.

License

The weights are under the LLAMA 2 Community License. Terms that affect derivatives:

  • Distributing the Llama Materials requires including the attribution notice: "Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved."
  • If your products or services had more than 700 million monthly active users in the month before the Llama 2 release date, you must request a license from Meta.
  • You may not use the Llama Materials or their output to improve any other large language model (excluding Llama 2 or derivative works thereof).
  • The Acceptable Use Policy applies to the model and to derivatives, and includes a prohibition on military and warfare applications.

Read the license text to confirm how it treats your fine-tuned adapter or merged model.

Memory to fine-tune Llama 2, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Llama-2-7b-hf6.7B100 GBRTX A5000 × 5 · $0.880/hr12.6 GBRTX A4000 · $0.167/hr3.2 GBRTX 4070 Super · $0.121/hr4.0 × 10^19 FLOPs
Llama-2-13b-chat-hf13.0B194 GBRTX A6000 × 5 · $1.81/hr24.2 GBRTX 4080 Super · $0.338/hr6.3 GBRTX 4070 Super · $0.121/hr7.8 × 10^19 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.