How to fine-tune GLM-4.5 (LoRA, SFT)
Fine-tuning GLM-4.5 and GLM-4.5-Air: base versus hybrid checkpoints, the tools the model card names, the MIT license and the chat template.
This guide covers adapting the GLM-4.5 family. For GPU counts and cost, see the sizing below and the training cost calculator. The fine-tuning overview covers the general workflow.
Which checkpoint to start from
The model card publishes both kinds of weights:
- Base models: GLM-4.5-Base and GLM-4.5-Air-Base, without instruction tuning. Choose these for continued pretraining or when you want your own instruction format.
- Hybrid reasoning models: GLM-4.5 and GLM-4.5-Air, which can think or answer directly. Choose these to keep chat and tool behavior.
- FP8 builds exist for serving. Train from the unquantized weights.
Tooling
The card states that fine-tuning runs under LLaMA Factory and under Swift, and its table lists LoRA and SFT strategies (plus RL). Strategies include LoRA. The card does not list LoRA target modules, so take them from the framework's GLM-4.5 template rather than guessing.
Chat template
Thinking is on by default in the hybrid models, and enable_thinking set to false turns it off. Train with the same setting you will serve with.
License
The card lists the license as MIT, described as allowing commercial use and secondary development. No extra limits on derivatives are stated.
MoE note
These are mixture-of-experts models with many more total parameters than active ones, so the frozen base still has to fit in memory even though few parameters are active per token.
Memory to fine-tune GLM-4.5, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| GLM-4.7-Flash | 31.2B | 465 GBRTX PRO 6000 × 5 · $6.88/hr | 58.2 GBA100 · $1.21/hr | 15.0 GBRTX A4000 · $0.167/hr | 1.9 × 10^20 FLOPs |
| GLM-4.5-Air | 110.5B | 1646 GBNo live fit | 206 GBRTX A6000 × 5 · $1.81/hr | 53.1 GBA100 · $1.21/hr | 6.6 × 10^20 FLOPs |
| GLM-4.6 | 356.8B | 5317 GBNo live fit | 665 GBRTX PRO 6000 × 7 · $9.63/hr | 171 GBRTX A5000 × 8 · $1.41/hr | 2.1 × 10^21 FLOPs |
| GLM-4.5 | 358.3B | 5340 GBNo live fit | 667 GBRTX PRO 6000 WS × 7 · $10.28/hr | 172 GBRTX A5000 × 8 · $1.41/hr | 2.2 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
Updated 2026-10-07.