How to fine-tune Qwen2 (LoRA, QLoRA)
Checkpoint choice, chat template, Transformers requirement and Apache 2.0 notes before fine-tuning Qwen2 with LoRA or QLoRA.
This guide covers adapting the Qwen2 family. Memory per method and GPU counts appear below, and the training cost calculator estimates runtime. For method background see fine-tuning, LoRA and QLoRA.
Checkpoint choice
- Start from the Instruct checkpoint to keep existing chat behavior and add a narrower skill.
- Start from the base checkpoint if you are training your own instruction behavior from scratch.
Chat template
Format training examples with the tokenizer's chat template (apply_chat_template, system, user and assistant roles) so training matches what the server applies later.
Requirements
transformers>=4.37.0. Older versions raiseKeyError: 'qwen2'.- The cards I opened do not document LoRA target modules, so use your trainer's defaults for this architecture.
- If you train for long inputs, remember the model's 131,072 token window relies on YaRN (rope scaling), and the card notes vLLM applies it statically.
License
The Qwen2-7B-Instruct card lists Apache 2.0. Sizes can differ, so check the license field on the exact checkpoint before you distribute a derivative.
Memory to fine-tune Qwen2, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Qwen2-0.5B | 494M | 7.4 GBRTX 4070 Super · $0.121/hr | 0.9 GBRTX 4070 Super · $0.121/hr | 0.2 GBRTX 4070 Super · $0.121/hr | 3.0 × 10^18 FLOPs |
| Qwen2-1.5B-Instruct | 1.5B | 23.0 GBRTX A5000 · $0.176/hr | 2.9 GBRTX 4070 Super · $0.121/hr | 0.7 GBRTX 4070 Super · $0.121/hr | 9.3 × 10^18 FLOPs |
| Qwen2-7B-Instruct | 7.6B | 113 GBRTX A5000 × 5 · $0.880/hr | 14.2 GBRTX A5000 · $0.176/hr | 3.7 GBRTX 4070 Super · $0.121/hr | 4.6 × 10^19 FLOPs |
| Qwen2-57B-A14B-Instruct | 57.4B | 855 GBNo live fit | 107 GBRTX A5000 × 5 · $0.880/hr | 27.6 GBRTX 4080 Super · $0.338/hr | 3.4 × 10^20 FLOPs |
| Qwen2-72B-Instruct | 72.7B | 1083 GBNo live fit | 135 GBRTX A5000 × 6 · $1.06/hr | 34.9 GBRTX A6000 · $0.363/hr | 4.4 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.