How to fine-tune Llama 4 (LoRA, QLoRA)
Checkpoint choice, license terms for derivatives and multimodal notes before fine-tuning Llama 4 Scout or Maverick with LoRA or QLoRA.
This guide covers adapting the Llama 4 family. For adapter methods see LoRA and QLoRA, and to size a job use the training cost calculator or the fine-tuning overview.
Checkpoint
The Scout card describes the Instruct checkpoint, the chat-tuned release. The card does not document a pretrained base checkpoint, so start from Instruct unless you confirm one on the family hub.
Tooling
The card requires Transformers 4.51.0 or higher, so use at least that in your training environment. It does not document LoRA target modules or MoE routing behavior, so confirm module names from the model config before you set them.
Multimodal
Llama 4 takes text and images as input. The card says it was tested with up to 5 images, so keep multi-image training examples within that or validate the results yourself.
License limits on derivatives
The Llama 4 Community License, as summarized on the card, requires:
- "Built with Llama" attribution on products or websites.
- "Llama" at the start of any derivative AI model name.
- A copy of the license with distributions.
- Compliance with the Acceptable Use Policy.
- A separate license from Meta if your products exceed 700 million monthly active users.
Read the full license text before shipping a fine-tuned model.
Memory to fine-tune Llama 4, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
No Llama 4 model of a size we can compute is in the catalog yet. See the Llama 4 model list for what is published.
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.