How to fine-tune Gemma 3 (LoRA, QLoRA)
Fine-tune Gemma 3 with LoRA or QLoRA: pt vs it checkpoints, the system prompt, bfloat16 training, and the Gemma Terms of Use for derivatives.
Gemma 3 spans 1B to 27B; the Gemma 3 hub lists sizes. For cost, see the training cost calculator and fine-tuning.
Checkpoint choice
The card offers pre-trained (PT) and instruction-tuned (IT) checkpoints. Start from PT for continued pre-training or a fully custom format, and from IT to keep chat behavior. Methods such as LoRA and QLoRA apply to both; the sources do not document target modules, so choose them from the layer names.
Data and template
- Train with the IT chat template. The card's examples use
systemanduserroles. - The launch post recommends
bfloat16, which also suits training. - The 1B model is text only. 4B and larger have a SigLIP vision tower at 896 x 896 input, so decide whether to freeze it or train only the language layers.
- The launch post mentions Unsloth integration for fine-tuning.
License limits on derivatives
Gemma 3 uses the Gemma Terms of Use, which you must accept on Hugging Face. Per the terms:
- A "Model Derivative" includes modifications and models trained to transfer Gemma's weights or patterns, but not its outputs.
- If you redistribute a derivative, you must pass on the use restrictions (Prohibited Use Policy), give recipients a copy of the terms, mark modified files, and include a notice file saying Gemma is provided under those terms.
Read the full terms before shipping a derivative.
Memory to fine-tune Gemma 3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| gemma-3-270m | 268M | 4.0 GBRTX 4070 Super · $0.121/hr | 0.5 GBRTX 4070 Super · $0.121/hr | 0.1 GBRTX 4070 Super · $0.121/hr | 1.6 × 10^18 FLOPs |
| gemma-3-1b-it | 1000M | 14.9 GBRTX A5000 · $0.176/hr | 1.9 GBRTX 4070 Super · $0.121/hr | 0.5 GBRTX 4070 Super · $0.121/hr | 6.0 × 10^18 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/google/gemma-3-27b-it
- https://huggingface.co/blog/gemma3
- https://ai.google.dev/gemma/terms
Updated 2026-10-07.