GLM models, generation by generation

Zhipu AI (Z.ai) GLM line from GLM-4 to GLM-5.3: sizes, MoE vs dense, context, licenses, and which generation to run today.

What the GLM line is

GLM is the general language model family from Zhipu AI, which publishes under the name Z.ai and the zai-org organization on Hugging Face. The line began with dense models of 9B to 32B parameters. Since GLM-4.5 it has been large mixture-of-experts (MoE) models aimed at coding and agent work, with MIT licensing for most releases. The newest flagship, GLM-5.3, is the exception and ships under its own license.

Generations in order

GLM-4 (April 2025 for the 0414 series, per Hugging Face repository dates). The GLM-4-0414 series includes GLM-4-9B-0414 and GLM-Z1-32B-0414, dense models under the MIT license. The 9B card describes it as aimed at resource-constrained scenarios, trained with the same reasoning-focused methods (extended reinforcement learning, math, code and logic data) as the larger GLM-4-32B-0414. The hub also holds GLM-4.1V-9B-Thinking, a vision-language model, and CodeGeeX4 9B.

GLM-4.5 (July 2025). The change here is architecture: GLM-4.5 is a MoE model with 355 billion total and 32 billion active parameters, and GLM-4.5-Air has 106 billion total and 12 billion active. Both have a 128K token context window, multi-token prediction layers, MIT licensing, and two modes: a thinking mode for reasoning and tool use, and a non-thinking mode for fast replies. The hub also holds the follow-ups GLM-4.6 (September 2025, context extended from 128K to 200K tokens according to its card) and GLM-4.7 (December 2025), plus GLM-4.7-Flash (January 2026), a 31 billion parameter 30B-A3B MoE under MIT. GLM-4.5V and GLM-4.6V-Flash are the vision variants.

GLM-5 (February 2026). GLM-5 scales to 744 billion total parameters with 40 billion active, up from 32 billion active in GLM-4.5, and its card says pre-training data grew from 23T to 28.5T tokens. It adopts DeepSeek Sparse Attention to cut deployment cost at long context, and it is MIT licensed. It is supported by vLLM, SGLang, KTransformers, Transformers and xLLM. The point releases GLM-5.1 (April 2026) and GLM-5.2 (June 2026) sit on the same hub.

GLM-5.3 (August 2026) keeps the same base architecture as GLM-5.2; its card attributes the gains entirely to post-training. Z.ai reports a 50 percent improvement over GLM-5.2 on its own in-house Z.ai Code Bench, a vendor-reported figure that has not been independently checked. The weights are about 753 billion parameters and are released under a custom glm-5.3 license rather than MIT, so read it before commercial use. The card lists BF16 and FP8 tensor types, and the hub carries both the BF16 and FP8 variants.

GLM-5.3-Flash (also August 2026) is a separate design rather than a trimmed GLM-5.3. It has 320 billion total and 18 billion active parameters, a hybrid of sparse and linear attention, and native image-text input, which its card calls the first natively multimodal model in the GLM-5 series. It is MIT licensed. Z.ai claims it outperforms GLM-5.2 at one-tenth the price (vendor-reported).

One pattern to notice across the line: active parameters grew from 12 billion (GLM-4.5-Air) to 32 billion (GLM-4.5) to 40 billion (GLM-5), and the Flash variants then pushed back toward lower active counts (3 billion in GLM-4.7-Flash, 18 billion in GLM-5.3-Flash). Total size, not active size, decides how much GPU memory a deployment needs.

Which generation to use today

  • A single mid-size GPU and a plain dense model: GLM-4-9B-0414. It is 9 billion parameters, MIT licensed and supported by Transformers, vLLM and SGLang.
  • Small MoE for agents and coding at low active compute: GLM-4.7-Flash (30B-A3B, MIT). Only about 3 billion parameters are active per token, though all 31 billion must be loaded.
  • Open weights you can use without license review: GLM-5 or GLM-5.3-Flash, both MIT. Choose GLM-5.3-Flash if you need image input.
  • The strongest GLM, if the license fits: GLM-5.3, a custom license at roughly 753 billion parameters, which means a multi-GPU deployment.
  • Older tooling pinned to GLM-4.5: GLM-4.5-Air (106B total, 12B active, 128K context) remains the lighter of the first MoE generation.

All of these can run on Aquanode GPUs.

GLM generations

Every GLM generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
GLM-469.4B to 32.6B21.0 GBRTX A5000$0.176/hr
GLM-4.51010.3B to 358.5B23.0 GBRTX A5000$0.176/hr
GLM-5221.2B to 753.9B359 GBRTX PRO 6000 × 4$5.50/hr

Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.

Run and fine-tune GLM

Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.

GLM compared with other models

Side-by-side VRAM, context length and license for models of a similar size.

Best models by task

GLM models appear on these ranked task pages.

Sources

Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.