How to run GLM-4.5 with vLLM and SGLang
Serving GLM-4.5 and GLM-4.5-Air with vLLM and SGLang: parser flags, the thinking toggle, FP8 variants and speculative decoding options.
This guide covers serving the GLM-4.5 family. Memory sizing and launch commands for each size appear below.
Engine support
The model card gives serving instructions for vLLM and SGLang, and a Transformers inference script. It does not state minimum versions for any of them, so use a recent release of each.
The card lists GLM-4.5 (355B total, 32B active) and GLM-4.5-Air (106B total, 12B active), each with an FP8 build (FP8). For FP8 on SGLang the card adds --disable-shared-experts-fusion.
Parsers and tool calling
Both engines use the same parser names:
--tool-call-parser glm45--reasoning-parser glm45--enable-auto-tool-choice(vLLM)
Thinking toggle
GLM-4.5 is a hybrid reasoning model and thinking is on by default. To turn it off per request, pass chat_template_kwargs with enable_thinking set to false through extra_body on vLLM or SGLang.
Speculative decoding
For SGLang the card shows EAGLE speculative decoding: --speculative-algorithm EAGLE, --speculative-num-steps 3, --speculative-eagle-topk 1 and --speculative-num-draft-tokens 4.
Context length
The card states a 128K token context. It documents no rope or YaRN override, so none is needed for the stated length.
Pitfalls
- On 8 H100 GPUs the full GLM-4.5 on vLLM needs
--cpu-offload-gb 16according to the card. - Reasoning and tool parsers are separate flags: set both or tool calls and thinking text will not be split out of the reply.
Launch commands by GLM-4.5 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
GLM-4.7-Flash (31.2B)
Needs about 69.8 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve zai-org/GLM-4.7-Flash --tensor-parallel-size 1SGLang
sglang serve --model-path zai-org/GLM-4.7-Flash --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/GLM-4.7-Flash-GGUFPrebuilt GGUF weights published at unsloth/GLM-4.7-Flash-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
GLM-4.5-Air (110.5B)
Needs about 247 GB of VRAM at BF16. Cheapest live fit: 6× RTX A6000 at $2.18/hr.
vLLM
vllm serve zai-org/GLM-4.5-Air --tensor-parallel-size 6SGLang
sglang serve --model-path zai-org/GLM-4.5-Air --tp 6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf ubergarm/GLM-4.5-Air-GGUFPrebuilt GGUF weights published at ubergarm/GLM-4.5-Air-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
GLM-4.6 (356.8B)
Needs about 797 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve zai-org/GLM-4.6SGLang
sglang serve --model-path zai-org/GLM-4.6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/GLM-4.6-GGUFPrebuilt GGUF weights published at unsloth/GLM-4.6-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
GLM-4.5 (358.3B)
Needs about 801 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve zai-org/GLM-4.5SGLang
sglang serve --model-path zai-org/GLM-4.5Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.