How to fine-tune Qwen3 (LoRA, QLoRA)
Which Qwen3 checkpoint to start from, chat template and thinking-mode handling, and the Apache 2.0 license terms for fine-tuning Qwen3.
This guide covers adapting models from the Qwen3 family. Memory per method and GPU counts are shown below. See also LoRA, QLoRA, the training cost calculator and the fine-tuning overview.
Which checkpoint to start from
- The post-trained Qwen3 checkpoints (for example Qwen3-32B) support both thinking and non-thinking modes. Start here if you want to keep chat and reasoning behavior and adapt it.
- Qwen publishes Base checkpoints for at least some sizes (Qwen3-8B-Base appears in the Qwen organization listing on Hugging Face). Start from a Base model for domain pretraining or when you will do your own instruction tuning.
Chat template and thinking mode
Train with the template that ships with the model, which exposes enable_thinking (default true) and the /think and /no_think soft switches. The card advises leaving thinking content out of multi-turn history. If your data has no reasoning traces, format examples consistently as non-thinking turns so the adapter does not learn mixed behavior.
Adapter tips
The model card does not document LoRA target modules. Target the attention and MLP projections as you would for other decoder models, and check module names in the checkpoint before you configure the adapter. The family includes both dense and Mixture-of-Experts sizes (the Qwen3MoE architecture is named in the llama.cpp docs), so confirm which you are tuning.
Use a recent Transformers release: versions below 4.51.0 raise a KeyError on Qwen3.
License
Qwen3-32B is released under Apache 2.0, which permits commercial use and derivative works under its standard terms. Check the license field on the specific checkpoint you use.
Memory to fine-tune Qwen3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Qwen3-0.6B | 752M | 11.2 GBRTX 4070 Super · $0.121/hr | 1.4 GBRTX 4070 Super · $0.121/hr | 0.4 GBRTX 4070 Super · $0.121/hr | 4.5 × 10^18 FLOPs |
| Qwen3-1.7B | 2.0B | 30.3 GBRTX 4080 Super · $0.338/hr | 3.8 GBRTX 4070 Super · $0.121/hr | 1.0 GBRTX 4070 Super · $0.121/hr | 1.2 × 10^19 FLOPs |
| Qwen3-4B | 4.0B | 59.9 GBA100 · $1.21/hr | 7.5 GBRTX 4070 Super · $0.121/hr | 1.9 GBRTX 4070 Super · $0.121/hr | 2.4 × 10^19 FLOPs |
| Qwen3-8B | 8.2B | 122 GBRTX A5000 × 6 · $1.06/hr | 15.3 GBRTX A4000 · $0.167/hr | 3.9 GBRTX 4070 Super · $0.121/hr | 4.9 × 10^19 FLOPs |
| Qwen3-14B | 14.8B | 220 GBRTX A6000 × 5 · $1.81/hr | 27.5 GBRTX 4080 Super · $0.338/hr | 7.1 GBRTX 4070 Super · $0.121/hr | 8.9 × 10^19 FLOPs |
| Qwen3-30B-A3B | 30.5B | 455 GBRTX PRO 6000 × 5 · $6.88/hr | 56.9 GBA100 · $1.21/hr | 14.7 GBRTX A4000 · $0.167/hr | 1.8 × 10^20 FLOPs |
| Qwen3-32B | 32.8B | 488 GBRTX PRO 6000 × 6 · $8.25/hr | 61.0 GBA100 · $1.21/hr | 15.7 GBRTX A4000 · $0.167/hr | 2.0 × 10^20 FLOPs |
| Qwen3-Coder-Next | 79.7B | 1187 GBNo live fit | 148 GBRTX A5000 × 7 · $1.23/hr | 38.3 GBRTX A6000 · $0.363/hr | 4.8 × 10^20 FLOPs |
| Qwen3-Next-80B-A3B-Instruct | 81.3B | 1212 GBNo live fit | 151 GBRTX A5000 × 7 · $1.23/hr | 39.1 GBRTX A6000 · $0.363/hr | 4.9 × 10^20 FLOPs |
| Qwen3-235B-A22B | 235.1B | 3503 GBNo live fit | 438 GBRTX PRO 6000 × 5 · $6.88/hr | 113 GBRTX A5000 × 5 · $0.880/hr | 1.4 × 10^21 FLOPs |
| Qwen3-Coder-480B-A35B-Instruct | 480.2B | 7155 GBNo live fit | 894 GBNo live fit | 231 GBRTX A6000 × 5 · $1.81/hr | 2.9 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.