How to fine-tune GLM-5 (LoRA, QLoRA)
MIT license, checkpoint choices, chat template and mixture-of-experts considerations when fine-tuning GLM-5.
This guide covers adapting the GLM-5 family. The card has no fine-tuning section, so this is limited to documented facts. See /fine-tuning and the training cost calculator for planning.
Which checkpoint
- The card describes one post-trained GLM-5 model. I did not find a separate base checkpoint on the card I read.
- An FP8 variant exists for serving (the vLLM recipe names GLM-5.1-FP8). For LoRA or QLoRA, start from the higher-precision weights and quantize afterward (FP8).
License
The card's license field is MIT, which permits modification and derivative works. Read the repository LICENSE file before redistributing.
Chat template
Use the model's own template through tokenizer.apply_chat_template() with standard role and content messages, so training examples match inference. Thinking is on by default and the template accepts enable_thinking. Decide whether your examples include reasoning content or train the non-thinking path.
Architecture notes
- Mixture-of-Experts: 744B parameters with 40B active.
- It integrates DeepSeek Sparse Attention.
- The card lists Transformers support, but it does not document LoRA target modules or an expert-routing recipe. Check that your trainer supports this architecture before planning a run.
Memory to fine-tune GLM-5, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| GLM-5.2 | 753.3B | 11225 GBNo live fit | 1403 GBNo live fit | 362 GBRTX A6000 × 8 · $2.90/hr | 4.5 × 10^21 FLOPs |
| GLM-5.1 | 753.9B | 11233 GBNo live fit | 1404 GBNo live fit | 362 GBRTX A6000 × 8 · $2.90/hr | 4.5 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/zai-org/GLM-5
- https://huggingface.co/zai-org/GLM-5/raw/main/README.md
- https://docs.vllm.ai/projects/recipes/en/latest/GLM/GLM5.html
Updated 2026-10-07.