How to fine-tune Gemma 4 (LoRA, QLoRA)
Fine-tune Gemma 4 with LoRA or QLoRA: pick base or -it checkpoints, freeze the vision and audio encoders, and check the Apache 2.0 license.
Gemma 4 comes in five variants; the Gemma 4 hub lists sizes. For cost, see the training cost calculator and fine-tuning.
Checkpoint choice
The card publishes pre-trained base checkpoints and instruction-tuned (-it) checkpoints. Start from the base for domain adaptation on raw text, or from -it to keep chat behavior and the native system role.
Method
- The Hugging Face launch post confirms PEFT and bitsandbytes compatibility, so LoRA and QLoRA both apply.
- TRL has full multimodal support, and Unsloth Studio is listed as a UI route.
- For multimodal models, freeze the vision and audio encoders and adapt only the language side, as the post advises.
- LoRA target modules are not documented in the sources, so pick them from the model's layer names.
Data and template
- Train with the same chat template you will serve with:
system,userandassistantroles. - If you want thinking behavior, the template uses
<|think|>in the system prompt and a<|channel>thoughtblock in the output. Keep your training format consistent with whichever mode you serve. - Put image content before text in multimodal samples.
MoE note
The 26B A4B variant is a mixture of experts with 3.8B active parameters. The sources give no MoE-specific LoRA guidance.
License
Gemma 4 is released under Apache 2.0 according to the model card, so derivatives carry no Gemma-specific terms beyond the Apache 2.0 conditions.
Memory to fine-tune Gemma 4, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
No Gemma 4 model of a size we can compute is in the catalog yet. See the Gemma 4 model list for what is published.
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.