How to fine-tune Gemma 4 (LoRA, QLoRA)

Fine-tune Gemma 4 with LoRA or QLoRA: pick base or -it checkpoints, freeze the vision and audio encoders, and check the Apache 2.0 license.

Gemma 4 comes in five variants; the Gemma 4 hub lists sizes. For cost, see the training cost calculator and fine-tuning.

Checkpoint choice

The card publishes pre-trained base checkpoints and instruction-tuned (-it) checkpoints. Start from the base for domain adaptation on raw text, or from -it to keep chat behavior and the native system role.

Method

  • The Hugging Face launch post confirms PEFT and bitsandbytes compatibility, so LoRA and QLoRA both apply.
  • TRL has full multimodal support, and Unsloth Studio is listed as a UI route.
  • For multimodal models, freeze the vision and audio encoders and adapt only the language side, as the post advises.
  • LoRA target modules are not documented in the sources, so pick them from the model's layer names.

Data and template

  • Train with the same chat template you will serve with: system, user and assistant roles.
  • If you want thinking behavior, the template uses <|think|> in the system prompt and a <|channel>thought block in the output. Keep your training format consistent with whichever mode you serve.
  • Put image content before text in multimodal samples.

MoE note

The 26B A4B variant is a mixture of experts with 3.8B active parameters. The sources give no MoE-specific LoRA guidance.

License

Gemma 4 is released under Apache 2.0 according to the model card, so derivatives carry no Gemma-specific terms beyond the Apache 2.0 conditions.

Memory to fine-tune Gemma 4, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

No Gemma 4 model of a size we can compute is in the catalog yet. See the Gemma 4 model list for what is published.

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.