How to run GLM-4.5 with vLLM and SGLang

Serving GLM-4.5 and GLM-4.5-Air with vLLM and SGLang: parser flags, the thinking toggle, FP8 variants and speculative decoding options.

This guide covers serving the GLM-4.5 family. Memory sizing and launch commands for each size appear below.

Engine support

The model card gives serving instructions for vLLM and SGLang, and a Transformers inference script. It does not state minimum versions for any of them, so use a recent release of each.

The card lists GLM-4.5 (355B total, 32B active) and GLM-4.5-Air (106B total, 12B active), each with an FP8 build (FP8). For FP8 on SGLang the card adds --disable-shared-experts-fusion.

Parsers and tool calling

Both engines use the same parser names:

  • --tool-call-parser glm45
  • --reasoning-parser glm45
  • --enable-auto-tool-choice (vLLM)

Thinking toggle

GLM-4.5 is a hybrid reasoning model and thinking is on by default. To turn it off per request, pass chat_template_kwargs with enable_thinking set to false through extra_body on vLLM or SGLang.

Speculative decoding

For SGLang the card shows EAGLE speculative decoding: --speculative-algorithm EAGLE, --speculative-num-steps 3, --speculative-eagle-topk 1 and --speculative-num-draft-tokens 4.

Context length

The card states a 128K token context. It documents no rope or YaRN override, so none is needed for the stated length.

Pitfalls

  • On 8 H100 GPUs the full GLM-4.5 on vLLM needs --cpu-offload-gb 16 according to the card.
  • Reasoning and tool parsers are separate flags: set both or tool calls and thinking text will not be split out of the reply.

Launch commands by GLM-4.5 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

GLM-4.7-Flash (31.2B)

Needs about 69.8 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve zai-org/GLM-4.7-Flash --tensor-parallel-size 1

SGLang

sglang serve --model-path zai-org/GLM-4.7-Flash --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run glm-4.7-flash

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/GLM-4.7-Flash-GGUF

Prebuilt GGUF weights published at unsloth/GLM-4.7-Flash-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

GLM-4.5-Air (110.5B)

Needs about 247 GB of VRAM at BF16. Cheapest live fit: 6× RTX A6000 at $2.18/hr.

vLLM

vllm serve zai-org/GLM-4.5-Air --tensor-parallel-size 6

SGLang

sglang serve --model-path zai-org/GLM-4.5-Air --tp 6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

llama.cpp

llama-server -hf ubergarm/GLM-4.5-Air-GGUF

Prebuilt GGUF weights published at ubergarm/GLM-4.5-Air-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

GLM-4.6 (356.8B)

Needs about 797 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve zai-org/GLM-4.6

SGLang

sglang serve --model-path zai-org/GLM-4.6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

llama.cpp

llama-server -hf unsloth/GLM-4.6-GGUF

Prebuilt GGUF weights published at unsloth/GLM-4.6-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

GLM-4.5 (358.3B)

Needs about 801 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve zai-org/GLM-4.5

SGLang

sglang serve --model-path zai-org/GLM-4.5

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.