How to fine-tune Llama 3.3 (LoRA, QLoRA)

Starting checkpoint, language coverage and Llama 3.3 Community License limits to check before fine-tuning Llama 3.3 70B with LoRA or QLoRA.

This guide covers adapting Llama 3.3. For adapter methods see LoRA and QLoRA. To size a run use the training cost calculator or the fine-tuning overview.

Checkpoint

The card describes the 70B Instruct model, tuned with supervised fine-tuning and RLHF, and says it is built on meta-llama/Llama-3.1-70B as its base. Start from Instruct to keep chat and tool behavior. The card does not publish a separate Llama 3.3 base checkpoint.

Chat template and tooling

Train with the tokenizer chat template shipped with the model, which the card says handles messages and tool formats. It requires Transformers 4.45.0 or newer, and documents 8-bit and 4-bit loading with bitsandbytes, the usual route for QLoRA. LoRA target modules are not documented on the card.

Languages

Eight languages are officially supported (English, German, French, Italian, Portuguese, Hindi, Spanish, Thai). The card allows fine-tuning for additional languages, provided you follow the license and acceptable use policy.

License limits on derivatives

Under the Llama 3.3 Community License as described on the card:

  • Show "Built with Llama" attribution.
  • Begin the name of any derivative model with "Llama".
  • Request permission from Meta if your monthly active users exceed 700 million.

Read the full license before distributing a fine-tuned model.

Memory to fine-tune Llama 3.3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Llama-3.3-70B-Instruct70.6B1051 GBNo live fit131 GBRTX A5000 × 6 · $1.06/hr33.9 GBRTX A6000 · $0.363/hr4.2 × 10^20 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.