How to fine-tune GLM-4.5 (LoRA, SFT)

Fine-tuning GLM-4.5 and GLM-4.5-Air: base versus hybrid checkpoints, the tools the model card names, the MIT license and the chat template.

This guide covers adapting the GLM-4.5 family. For GPU counts and cost, see the sizing below and the training cost calculator. The fine-tuning overview covers the general workflow.

Which checkpoint to start from

The model card publishes both kinds of weights:

  • Base models: GLM-4.5-Base and GLM-4.5-Air-Base, without instruction tuning. Choose these for continued pretraining or when you want your own instruction format.
  • Hybrid reasoning models: GLM-4.5 and GLM-4.5-Air, which can think or answer directly. Choose these to keep chat and tool behavior.
  • FP8 builds exist for serving. Train from the unquantized weights.

Tooling

The card states that fine-tuning runs under LLaMA Factory and under Swift, and its table lists LoRA and SFT strategies (plus RL). Strategies include LoRA. The card does not list LoRA target modules, so take them from the framework's GLM-4.5 template rather than guessing.

Chat template

Thinking is on by default in the hybrid models, and enable_thinking set to false turns it off. Train with the same setting you will serve with.

License

The card lists the license as MIT, described as allowing commercial use and secondary development. No extra limits on derivatives are stated.

MoE note

These are mixture-of-experts models with many more total parameters than active ones, so the frozen base still has to fit in memory even though few parameters are active per token.

Memory to fine-tune GLM-4.5, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
GLM-4.7-Flash31.2B465 GBRTX PRO 6000 × 5 · $6.88/hr58.2 GBA100 · $1.21/hr15.0 GBRTX A4000 · $0.167/hr1.9 × 10^20 FLOPs
GLM-4.5-Air110.5B1646 GBNo live fit206 GBRTX A6000 × 5 · $1.81/hr53.1 GBA100 · $1.21/hr6.6 × 10^20 FLOPs
GLM-4.6356.8B5317 GBNo live fit665 GBRTX PRO 6000 × 7 · $9.63/hr171 GBRTX A5000 × 8 · $1.41/hr2.1 × 10^21 FLOPs
GLM-4.5358.3B5340 GBNo live fit667 GBRTX PRO 6000 WS × 7 · $10.28/hr172 GBRTX A5000 × 8 · $1.41/hr2.2 × 10^21 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.