How to fine-tune Gemma 2 (LoRA, QLoRA)
Fine-tune Gemma 2 with LoRA or QLoRA: pt vs it checkpoints, the turn-based template, and Gemma Terms of Use rules for derivatives.
The Gemma 2 card gives no official fine-tuning guidance, so this page covers only what is stated. See the Gemma 2 hub for sizes, and fine-tuning with the training cost calculator for planning.
Checkpoint choice
Pre-trained (PT) and instruction-tuned (IT) checkpoints both exist. Start from PT for continued pre-training or a custom format, and from IT to keep chat behavior. LoRA and QLoRA are standard options; target modules are not documented in the sources.
Template and dtype
- Train with the IT template: turns wrapped in
<start_of_turn>and<end_of_turn>, rolesuserandmodel. - The card describes no system role, so put instructions in the user turn.
- Weights are native
bfloat16, and upcasting tofloat32adds no precision, so train inbfloat16.
License limits on derivatives
Gemma 2 is under the Gemma Terms of Use. Per the terms:
- A "Model Derivative" includes modifications and models trained to transfer Gemma's weights or patterns, but not its outputs.
- Redistributing a derivative requires passing on the use restrictions (Prohibited Use Policy), giving recipients a copy of the terms, marking modified files, and including a notice file stating Gemma is provided under the terms.
Read the full terms before shipping a derivative.
Memory to fine-tune Gemma 2, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| gemma-2b-it | 2.5B | 37.3 GBRTX A6000 · $0.363/hr | 4.7 GBRTX 4070 Super · $0.121/hr | 1.2 GBRTX 4070 Super · $0.121/hr | 1.5 × 10^19 FLOPs |
| gemma-2-2b-it | 2.6B | 39.0 GBRTX A6000 · $0.363/hr | 4.9 GBRTX 4070 Super · $0.121/hr | 1.3 GBRTX 4070 Super · $0.121/hr | 1.6 × 10^19 FLOPs |
| gemma-2-9b-it | 9.2B | 138 GBRTX A5000 × 6 · $1.06/hr | 17.2 GBRTX A5000 · $0.176/hr | 4.4 GBRTX 4070 Super · $0.121/hr | 5.5 × 10^19 FLOPs |
| gemma-2-27b-it | 27.2B | 406 GBRTX PRO 6000 × 5 · $6.88/hr | 50.7 GBA100 · $1.21/hr | 13.1 GBRTX A4000 · $0.167/hr | 1.6 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.