How to fine-tune SmolLM3 (LoRA, QLoRA)
Fine-tune SmolLM3-3B with LoRA or QLoRA: base versus instruct, the /think and /no_think template, and its Apache 2.0 license.
The SmolLM3 hub has size details. At 3B parameters, LoRA or QLoRA fits easily on a single GPU. Estimate cost with the training cost calculator or see the fine-tuning overview.
Which checkpoint
SmolLM3-3Bis the instruction-tuned model the card recommends. Its post-training included midtraining on reasoning data, supervised fine-tuning and Anchored Preference Optimization.SmolLM3-3B-Baseis the pretrained model, for when you want to run your own post-training.- The authors also publish intermediate pretraining checkpoints in a separate repository.
The card gives no LoRA target modules, so use your library's defaults for this architecture.
Chat template and reasoning
The template switches reasoning with /think and /no_think in the system prompt, or enable_thinking=False in apply_chat_template. If you fine-tune the instruct model and want both modes to keep working, include examples of each, with the matching system flag, and render them through the tokenizer's template. Training only on one style risks shifting the default behavior.
Context
The default context is 65,536 tokens, with YaRN extension to 131,072 documented in the card. Keep training sequence lengths within what you intend to serve. See RoPE scaling.
License
The model card lists Apache 2.0, so there are no naming or user-count conditions beyond the standard Apache terms (attribution and notices).
Memory to fine-tune SmolLM3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| SmolLM3-3B | 3.1B | 45.8 GBRTX A6000 · $0.363/hr | 5.7 GBRTX 4070 Super · $0.121/hr | 1.5 GBRTX 4070 Super · $0.121/hr | 1.8 × 10^19 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.