How to fine-tune Granite 3.3 (LoRA, QLoRA)

Base versus instruct checkpoints, chat template and license notes for fine-tuning IBM Granite 3.3 with LoRA or QLoRA.

This guide covers adapting the Granite 3 family with LoRA or QLoRA. Estimate cost with the training cost calculator and see fine-tuning for the overview.

Which checkpoint

  • Granite 3.3 8B Base is a decoder-only model with 8.1B parameters trained on 12 trillion tokens, with a 128K context window and Fill-in-the-Middle code support. The card says base models "can serve as baseline to create specialized models for specific application scenarios."
  • Granite 3.3 8B Instruct already follows instructions and supports a thinking mode. Start from it when you want to keep chat and reasoning behaviour and only add a domain.
  • Start from Base when you plan full instruction tuning of your own.

Chat template

Train with the tokenizer's built-in template (apply_chat_template, role and content messages). If you want reasoning, keep the instruct format where thoughts sit in <think></think> and answers in <response></response>, and apply the same tags consistently in your data.

Target modules

The cards do not list recommended LoRA target modules. Use the attention and MLP projection names from the checkpoint and confirm them against the loaded model.

License

Both checkpoints are listed as Apache 2.0, which permits commercial use and derivative models. Check the repo's license file for attribution wording.

Languages

The cards say users may fine-tune Granite 3.3 for languages beyond the 12 supported.

Memory to fine-tune Granite 3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
granite-3.0-1b-a400m-instruct1.3B19.9 GBRTX A5000 · $0.176/hr2.5 GBRTX 4070 Super · $0.121/hr0.6 GBRTX 4070 Super · $0.121/hr8.0 × 10^18 FLOPs
granite-3.3-2b-instruct2.5B37.8 GBRTX A6000 · $0.363/hr4.7 GBRTX 4070 Super · $0.121/hr1.2 GBRTX 4070 Super · $0.121/hr1.5 × 10^19 FLOPs
granite-3.0-8b-instruct8.2B122 GBRTX A5000 × 6 · $1.06/hr15.2 GBRTX A4000 · $0.167/hr3.9 GBRTX 4070 Super · $0.121/hr4.9 × 10^19 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.