How to run GLM-5 with vLLM and SGLang

Minimum vLLM and SGLang versions, reasoning and tool-call parsers, thinking toggle and speculative decoding flags for GLM-5.

This guide covers serving the GLM-5 family. Sizing and launch commands per model appear below.

Engine support

The GLM-5 card lists these minimum versions:

  • vLLM 0.19.0 or newer.
  • SGLang 0.5.10 or newer.
  • Transformers support is also listed.

The card does not mention llama.cpp or Ollama.

Parsers and flags

Both engines use the same parser names on the card:

  • --tool-call-parser glm47 (with --enable-auto-tool-choice in vLLM)
  • --reasoning-parser glm45
  • Speculative decoding: --speculative-config.method mtp in vLLM, --speculative-algorithm EAGLE in SGLang.

Thinking toggle

The vLLM recipe says thinking is enabled by default and can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}. Messages go through the standard chat template (apply_chat_template).

Model facts

GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active) and integrates DeepSeek Sparse Attention. The license field is MIT.

Pitfalls

  • FP8 (FP8) checkpoints need DeepGEMM, installed with a script given in the vLLM recipe.
  • Disable prefix caching when benchmarking, or throughput looks inflated.
  • The recipe says tool calling together with MTP needs the latest vLLM main branch (stated there for GLM-5.1).

Launch commands by GLM-5 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

GLM-5.2 (753.3B)

Needs about 1684 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve zai-org/GLM-5.2

SGLang

sglang serve --model-path zai-org/GLM-5.2

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run glm-5.2

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/GLM-5.2-GGUF

Prebuilt GGUF weights published at unsloth/GLM-5.2-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

GLM-5.1 (753.9B)

Needs about 1685 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve zai-org/GLM-5.1

SGLang

sglang serve --model-path zai-org/GLM-5.1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

llama.cpp

llama-server -hf unsloth/GLM-5.1-GGUF

Prebuilt GGUF weights published at unsloth/GLM-5.1-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.