How to fine-tune DeepSeek-R1 (LoRA, QLoRA)
Which DeepSeek-R1 distilled checkpoint to start from, the MIT license and its base-model licenses, and prompt format for tuning.
This guide covers adapting the DeepSeek-R1 family. The card has no fine-tuning instructions, so this is limited to what it documents. See /fine-tuning and the training cost calculator for planning.
Which checkpoint
For LoRA or QLoRA, the practical starting points are the six distilled models, not the full R1:
- Qwen-based: 1.5B, 7B, 14B, 32B, fine-tuned from Qwen2.5 models using 800k samples curated with DeepSeek-R1.
- Llama-based: 8B (from Llama3.1-8B-Base) and 70B (from Llama3.3-70B-Instruct).
License
- The card says the code repository and weights are under the MIT License, supporting commercial use, modification and derivative works.
- The distilled models inherit their base-model terms: Apache 2.0 for the Qwen-based ones, the llama3.1 license for the 8B and the llama3.3 license for the 70B. Read those licenses before releasing a derivative.
Prompt format
Follow the usage recommendations when building training data:
- No system prompt, with all instructions in the user turn.
- Reasoning goes first, and the card recommends responses begin with
<think>\n. - Keep reasoning traces in the same style as the model's own output so tuning does not fight its behavior.
What the card does not say
LoRA target modules and MoE routing guidance for the full R1 are not documented on the card.
Memory to fine-tune DeepSeek R1, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| DeepSeek-R1-0528-Qwen3-8B | 8.2B | 122 GBRTX A5000 × 6 · $1.06/hr | 15.3 GBRTX A4000 · $0.167/hr | 3.9 GBRTX 4070 Super · $0.121/hr | 4.9 × 10^19 FLOPs |
| DeepSeek-R1 | 684.5B | 10200 GBNo live fit | 1275 GBNo live fit | 329 GBRTX A6000 × 7 · $2.54/hr | 4.1 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/deepseek-ai/DeepSeek-R1
- https://huggingface.co/deepseek-ai/DeepSeek-R1/raw/main/README.md
Updated 2026-10-07.