How to fine-tune Llama 3.2 (LoRA, QLoRA)

Base vs instruct checkpoints, quantized variants and Llama 3.2 Community License terms, including the EU vision limit, before fine-tuning.

This guide covers adapting Llama 3.2. For adapter methods see LoRA and QLoRA. To size a run use the training cost calculator or the fine-tuning overview.

Base or instruct

The 1B card describes both a pretrained base (trained on 9 trillion tokens) and an instruction-tuned model built with supervised fine-tuning and RLHF for dialogue and agentic use. The 11B Vision card likewise says a base pretrained model exists. Use Instruct to keep chat behavior, and the base model for continued pretraining or fully custom formats.

Chat template and tooling

Use the tokenizer or processor chat template. For the vision model, processor.apply_chat_template() formats messages that include images, with Transformers 4.45.0 or newer. LoRA target modules are not documented on the cards.

Quantized releases

The 1B card lists SpinQuant (4-bit groupwise weights with 8-bit dynamic activations) and QLoRA-trained quantized variants. These are published inference checkpoints, so do not assume they are suitable training starting points.

License limits on derivatives

  • The Llama 3.2 Community License asks for attribution and compliance with the Acceptable Use Policy.
  • A separate license from Meta is required above 700 million monthly active users.
  • For the multimodal models, the card states the Section 1(a) rights are not granted to individuals or companies based in the European Union. That applies to a fine-tuned vision model you build.

Read the full license text before distributing.

Memory to fine-tune Llama 3.2, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Llama-3.2-1B-Instruct1.2B18.4 GBRTX A5000 · $0.176/hr2.3 GBRTX 4070 Super · $0.121/hr0.6 GBRTX 4070 Super · $0.121/hr7.4 × 10^18 FLOPs
Llama-3.2-3B-Instruct3.2B47.9 GBRTX A6000 · $0.363/hr6.0 GBRTX 4070 Super · $0.121/hr1.5 GBRTX 4070 Super · $0.121/hr1.9 × 10^19 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.