How to run GLM-5 with vLLM and SGLang
Minimum vLLM and SGLang versions, reasoning and tool-call parsers, thinking toggle and speculative decoding flags for GLM-5.
This guide covers serving the GLM-5 family. Sizing and launch commands per model appear below.
Engine support
The GLM-5 card lists these minimum versions:
- vLLM 0.19.0 or newer.
- SGLang 0.5.10 or newer.
- Transformers support is also listed.
The card does not mention llama.cpp or Ollama.
Parsers and flags
Both engines use the same parser names on the card:
--tool-call-parser glm47(with--enable-auto-tool-choicein vLLM)--reasoning-parser glm45- Speculative decoding:
--speculative-config.method mtpin vLLM,--speculative-algorithm EAGLEin SGLang.
Thinking toggle
The vLLM recipe says thinking is enabled by default and can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}. Messages go through the standard chat template (apply_chat_template).
Model facts
GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active) and integrates DeepSeek Sparse Attention. The license field is MIT.
Pitfalls
- FP8 (FP8) checkpoints need DeepGEMM, installed with a script given in the vLLM recipe.
- Disable prefix caching when benchmarking, or throughput looks inflated.
- The recipe says tool calling together with MTP needs the latest vLLM main branch (stated there for GLM-5.1).
Launch commands by GLM-5 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
GLM-5.2 (753.3B)
Needs about 1684 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve zai-org/GLM-5.2SGLang
sglang serve --model-path zai-org/GLM-5.2Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/GLM-5.2-GGUFPrebuilt GGUF weights published at unsloth/GLM-5.2-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
GLM-5.1 (753.9B)
Needs about 1685 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve zai-org/GLM-5.1SGLang
sglang serve --model-path zai-org/GLM-5.1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/GLM-5.1-GGUFPrebuilt GGUF weights published at unsloth/GLM-5.1-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/zai-org/GLM-5
- https://huggingface.co/zai-org/GLM-5/raw/main/README.md
- https://docs.vllm.ai/projects/recipes/en/latest/GLM/GLM5.html
Updated 2026-10-07.