How to fine-tune Phi-4 (LoRA, QLoRA)

Fine-tune Phi-4 with LoRA or QLoRA: the single aligned checkpoint, its chat template, and the MIT license for derivatives.

See the Phi-4 hub for sizes. For planning, use the training cost calculator and fine-tuning.

Checkpoint choice

The card offers one variant: an instruction-aligned model produced with supervised fine-tuning and direct preference optimization. There is no separate base checkpoint listed, so adapters start from this model. LoRA and QLoRA are standard routes, and the card lists many community fine-tunes and adapters. It documents no target modules, so choose them from the layer names.

Data and template

  • Train with the model's chat format: system, user and assistant turns using <|im_start|>, <|im_sep|> and <|im_end|>.
  • The context length is 16K tokens, so keep training sequences within it.
  • Training data was mostly English and Python. Expect to need more data to move non-English or other-language behavior.

License

The model card lists the MIT license, which permits modification and redistribution of derivatives subject to MIT's notice requirement. Check the LICENSE file on the repository for the exact text.

Memory to fine-tune Phi-4, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Phi-4-mini-instruct3.8B57.2 GBA100 · $1.21/hr7.1 GBRTX 4070 Super · $0.121/hr1.8 GBRTX 4070 Super · $0.121/hr2.3 × 10^19 FLOPs
phi-414.7B218 GBRTX A6000 × 5 · $1.81/hr27.3 GBRTX 4080 Super · $0.338/hr7.0 GBRTX 4070 Super · $0.121/hr8.8 × 10^19 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.