How to fine-tune Falcon3 (LoRA, QLoRA)
Fine-tune Falcon3 with LoRA or QLoRA: base versus instruct checkpoints, chat template, and the TII Falcon License 2.0 terms for derivatives.
See the Falcon3 hub for sizes. LoRA and QLoRA are the usual routes on a single GPU. Estimate cost with the training cost calculator or read the fine-tuning overview.
Which checkpoint
- The Base checkpoint is, in the card's words, a raw pretrained model that should be further fine-tuned for most use cases. Start here for full control of behavior.
- The Instruct checkpoint was post-trained on 1.2 million samples of STEM, conversation, code, safety and function-call data. Start here to adapt a model that already follows instructions.
The cards document no LoRA target modules; use your library's defaults for this architecture (28 decoder blocks with grouped query attention).
Chat template
Format Instruct training data with the tokenizer's apply_chat_template using system and user roles, so training matches inference. The card describes no reasoning toggle.
Languages and context
The cards list English, French, Spanish and Portuguese, and a 32K context window.
License limits on derivatives
The TII terms for Falcon 3 and newer models say:
- Include prominently in any public statement about a derivative: "[name of derivative] is built using artificial intelligence technology from the Technology Innovation Institute" (TII allows reasonable adjustments on request).
- Your use must comply with TII's Acceptable Use Policy, and distribution agreements must carry use-based restrictions that incorporate it.
- Give recipients a copy of the license, mark modified files, and retain existing notices.
- Copyright and patent licenses are royalty free.
Read the full license before shipping a derivative commercially.
Memory to fine-tune Falcon3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Falcon3-1B-Instruct | 1.7B | 24.9 GBRTX 4080 Super · $0.338/hr | 3.1 GBRTX 4070 Super · $0.121/hr | 0.8 GBRTX 4070 Super · $0.121/hr | 1.0 × 10^19 FLOPs |
| Falcon3-7B-Instruct | 7.5B | 111 GBRTX A5000 × 5 · $0.880/hr | 13.9 GBRTX A4000 · $0.167/hr | 3.6 GBRTX 4070 Super · $0.121/hr | 4.5 × 10^19 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/tiiuae/Falcon3-7B-Instruct
- https://huggingface.co/tiiuae/Falcon3-7B-Base
- https://falconllm.tii.ae/falcon-terms-and-conditions.html
Updated 2026-10-07.