How to fine-tune Granite 4.0 (LoRA, QLoRA)
Base versus instruct checkpoints, chat template, license and hybrid-architecture notes for fine-tuning IBM Granite 4.0 with LoRA or QLoRA.
This guide covers adapting the Granite 4.0 family with LoRA or QLoRA. Estimate cost with the training cost calculator and see fine-tuning for the overview.
Which checkpoint
- The
-basecheckpoints are pretrained decoder-only models. The base card says they "can serve as baseline to create specialized models for specific application scenarios" and support Fill-in-the-Middle code completion. - The instruct checkpoints are the right start for chat, tool use and RAG-style behaviour, since they already carry the Granite chat template.
- Sizes: Micro Dense (3B), H Micro Dense (3B), H Tiny MoE (7B), H Small MoE (32B, 9B active).
Chat template
Train with the model's own template, which wraps turns in <|start_of_role|> and <|end_of_role|> tokens. Tool calls are JSON inside <tool_call> tags following the OpenAI function schema, so format tool-use data the same way. The instruct card notes a default system prompt added in October 2025.
Architecture notes
H models mix 36 Mamba2 layers with 4 attention layers, and H Small and H Tiny use mixture-of-experts. The cards do not document recommended LoRA target modules, so choose them from the module names in the checkpoint and verify your trainer supports the Mamba2 layers. The cards give no guidance on MoE routing during fine-tuning.
License
The checkpoints are listed as Apache 2.0, which allows commercial use and derivatives. Read the license file shipped with the repo for attribution terms.
Languages
The cards say users may fine-tune for languages beyond the 12 supported ones.
Memory to fine-tune Granite 4, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| granite-4.0-350m | 352M | 5.3 GBRTX 4070 Super · $0.121/hr | 0.7 GBRTX 4070 Super · $0.121/hr | 0.2 GBRTX 4070 Super · $0.121/hr | 2.1 × 10^18 FLOPs |
| granite-4.0-1b | 1.6B | 24.3 GBRTX 4080 Super · $0.338/hr | 3.0 GBRTX 4070 Super · $0.121/hr | 0.8 GBRTX 4070 Super · $0.121/hr | 9.8 × 10^18 FLOPs |
| granite-4.0-h-micro | 3.2B | 47.6 GBRTX A6000 · $0.363/hr | 5.9 GBRTX 4070 Super · $0.121/hr | 1.5 GBRTX 4070 Super · $0.121/hr | 1.9 × 10^19 FLOPs |
| granite-4.1-3b | 3.4B | 50.7 GBA100 · $1.21/hr | 6.3 GBRTX 4070 Super · $0.121/hr | 1.6 GBRTX 4070 Super · $0.121/hr | 2.0 × 10^19 FLOPs |
| granite-4.2-3b | 3.7B | 54.5 GBA100 · $1.21/hr | 6.8 GBRTX 4070 Super · $0.121/hr | 1.8 GBRTX 4070 Super · $0.121/hr | 2.2 × 10^19 FLOPs |
| granite-4.0-h-tiny | 6.9B | 103 GBRTX A5000 × 5 · $0.880/hr | 12.9 GBRTX A4000 · $0.167/hr | 3.3 GBRTX 4070 Super · $0.121/hr | 4.2 × 10^19 FLOPs |
| granite-4.1-8b | 8.8B | 131 GBRTX A5000 × 6 · $1.06/hr | 16.4 GBRTX A5000 · $0.176/hr | 4.2 GBRTX 4070 Super · $0.121/hr | 5.3 × 10^19 FLOPs |
| granite-4.1-30b | 28.9B | 430 GBRTX PRO 6000 × 5 · $6.88/hr | 53.8 GBA100 · $1.21/hr | 13.9 GBRTX A4000 · $0.167/hr | 1.7 × 10^20 FLOPs |
| granite-4.2-30b | 29.3B | 436 GBRTX PRO 6000 × 5 · $6.88/hr | 54.5 GBA100 · $1.21/hr | 14.1 GBRTX A4000 · $0.167/hr | 1.8 × 10^20 FLOPs |
| granite-4.0-h-small | 32.2B | 480 GBA100 × 6 · $7.27/hr | 60.0 GBA100 · $1.21/hr | 15.5 GBRTX A4000 · $0.167/hr | 1.9 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/ibm-granite/granite-4.0-h-small
- https://huggingface.co/ibm-granite/granite-4.0-h-small-base
Updated 2026-10-07.