How to fine-tune DeepSeek-V4 (LoRA, QLoRA)
Base checkpoints, license, custom chat encoding and MoE considerations when fine-tuning DeepSeek-V4 Pro and Flash.
This guide covers adapting the DeepSeek-V4 family. Documentation for fine-tuning is thin, so treat this as a starting checklist. See /fine-tuning and the training cost calculator for planning.
Which checkpoint
- A Base repository exists for Flash (
DeepSeek-V4-Flash-Base). - The Flash-Base page does not define "Base" or state a license. It lists tensor types BF16, I64, F32 and F8_E4M3.
- The post-trained models are DeepSeek-V4-Pro and DeepSeek-V4-Flash.
License
The DeepSeek-V4-Pro card says: "This repository and the model weights are licensed under the MIT License." I did not find a license statement on the Flash-Base page, so check the repository's LICENSE file before relying on it for derivatives.
Chat format
The model uses a custom encoding instead of a Jinja template. The card points to its encoding folder for converting messages into model input. Format training data with that code so the examples match what the model sees at inference. The card describes Non-think, Think High and Think Max modes, so decide which mode your training examples represent.
Architecture notes
- Both models are Mixture-of-Experts: Pro has 1.6T total and 49B activated parameters, Flash has 284B total and 13B activated.
- Expert weights ship as FP4 and most other weights as FP8, so LoRA and QLoRA tooling must handle these formats. The cards do not document target modules or a supported training recipe.
Memory to fine-tune DeepSeek V4, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 | 304.2B | 4533 GBNo live fit | 567 GBRTX PRO 6000 × 6 · $8.25/hr | 146 GBRTX A5000 × 7 · $1.23/hr | 1.8 × 10^21 FLOPs |
| DeepSeek-V4-Pro-0813 | 1650.5B | 24594 GBNo live fit | 3074 GBNo live fit | 793 GBNo live fit | 9.9 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base
Updated 2026-10-07.