How to fine-tune Qwen2.5 (LoRA, QLoRA)
Which Qwen2.5 checkpoint to start from, chat template and Apache 2.0 licensing notes before fine-tuning with LoRA or QLoRA.
This guide covers adapting the Qwen2.5 family. Memory per method and GPU counts appear below, and the training cost calculator estimates runtime. For method background see fine-tuning, LoRA and QLoRA.
Base or instruct
- Start from the base checkpoint when you have a large instruction dataset or want full control of behavior. The Qwen2.5-7B card says not to use base models for conversations and recommends post-training such as SFT, RLHF or continued pretraining.
- Start from the Instruct checkpoint when you want to keep existing chat behavior and add a narrower skill.
Chat template
Instruct checkpoints use a chat template with system, user and assistant roles, applied through apply_chat_template(). Format your training data with the same template you will serve with, so the tokens match at inference.
Requirements
transformers>=4.37.0, otherwise loading fails withKeyError: 'qwen2'.- The 7B model uses grouped-query attention (28 query heads, 4 key/value heads), per the card.
- LoRA target modules are not documented on the cards I opened, so pick them from your trainer's defaults.
License
The 7B base and Instruct cards list Apache 2.0. Not every Qwen2.5 size may carry the same license, so confirm the license field on the exact checkpoint you adapt before distributing derivatives.
Memory to fine-tune Qwen2.5, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Qwen2.5-0.5B-Instruct | 494M | 7.4 GBRTX 4070 Super · $0.121/hr | 0.9 GBRTX 4070 Super · $0.121/hr | 0.2 GBRTX 4070 Super · $0.121/hr | 3.0 × 10^18 FLOPs |
| Qwen2.5-1.5B-Instruct | 1.5B | 23.0 GBRTX A5000 · $0.176/hr | 2.9 GBRTX 4070 Super · $0.121/hr | 0.7 GBRTX 4070 Super · $0.121/hr | 9.3 × 10^18 FLOPs |
| Qwen2.5-3B-Instruct | 3.1B | 46.0 GBRTX A6000 · $0.363/hr | 5.7 GBRTX 4070 Super · $0.121/hr | 1.5 GBRTX 4070 Super · $0.121/hr | 1.9 × 10^19 FLOPs |
| Qwen2.5-7B-Instruct | 7.6B | 113 GBRTX A5000 × 5 · $0.880/hr | 14.2 GBRTX A4000 · $0.167/hr | 3.7 GBRTX 4070 Super · $0.121/hr | 4.6 × 10^19 FLOPs |
| Qwen2.5-14B-Instruct | 14.8B | 220 GBRTX A6000 × 5 · $1.81/hr | 27.5 GBRTX 4080 Super · $0.338/hr | 7.1 GBRTX 4070 Super · $0.121/hr | 8.9 × 10^19 FLOPs |
| Qwen2.5-32B-Instruct | 32.8B | 488 GBRTX PRO 6000 × 6 · $8.25/hr | 61.0 GBA100 · $1.21/hr | 15.7 GBRTX A4000 · $0.167/hr | 2.0 × 10^20 FLOPs |
| Qwen2.5-72B-Instruct | 72.7B | 1083 GBNo live fit | 135 GBRTX A5000 × 6 · $1.06/hr | 34.9 GBRTX A6000 · $0.363/hr | 4.4 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.