How to fine-tune Gemma 3 (LoRA, QLoRA)

Fine-tune Gemma 3 with LoRA or QLoRA: pt vs it checkpoints, the system prompt, bfloat16 training, and the Gemma Terms of Use for derivatives.

Gemma 3 spans 1B to 27B; the Gemma 3 hub lists sizes. For cost, see the training cost calculator and fine-tuning.

Checkpoint choice

The card offers pre-trained (PT) and instruction-tuned (IT) checkpoints. Start from PT for continued pre-training or a fully custom format, and from IT to keep chat behavior. Methods such as LoRA and QLoRA apply to both; the sources do not document target modules, so choose them from the layer names.

Data and template

  • Train with the IT chat template. The card's examples use system and user roles.
  • The launch post recommends bfloat16, which also suits training.
  • The 1B model is text only. 4B and larger have a SigLIP vision tower at 896 x 896 input, so decide whether to freeze it or train only the language layers.
  • The launch post mentions Unsloth integration for fine-tuning.

License limits on derivatives

Gemma 3 uses the Gemma Terms of Use, which you must accept on Hugging Face. Per the terms:

  • A "Model Derivative" includes modifications and models trained to transfer Gemma's weights or patterns, but not its outputs.
  • If you redistribute a derivative, you must pass on the use restrictions (Prohibited Use Policy), give recipients a copy of the terms, mark modified files, and include a notice file saying Gemma is provided under those terms.

Read the full terms before shipping a derivative.

Memory to fine-tune Gemma 3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
gemma-3-270m268M4.0 GBRTX 4070 Super · $0.121/hr0.5 GBRTX 4070 Super · $0.121/hr0.1 GBRTX 4070 Super · $0.121/hr1.6 × 10^18 FLOPs
gemma-3-1b-it1000M14.9 GBRTX A5000 · $0.176/hr1.9 GBRTX 4070 Super · $0.121/hr0.5 GBRTX 4070 Super · $0.121/hr6.0 × 10^18 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.