How to run Hy4 with vLLM and SGLang
Serving Tencent Hy4-preview with vLLM and SGLang: prebuilt images, parser flags, MTP speculative decoding and the reasoning_effort toggle.
This guide covers serving the Hy4 family. The checkpoint is named Hy4-preview, with an FP8 build (FP8). Memory sizing and launch commands appear below.
Engine support
The card documents vLLM and SGLang only, each through a prebuilt image: vllm/vllm-openai:hy4-preview and lmsysorg/sglang:hy4-preview (the latter for x86 and ARM). It states no minimum version, so use those images rather than a stock release. It does not cover llama.cpp or Ollama.
vLLM flags
--tensor-parallel-size 8--speculative-configwith methodmtpand 3 speculative tokens--attention-backend FLASHMLA_SPARSE--tool-call-parser hy_v4--reasoning-parser hy_v4
SGLang flags
--tp-size 8--reasoning-parser auto--speculative-algorithm NEXTNwith--speculative-num-steps 3
Reasoning toggle
Reasoning defaults to high. For direct answers pass chat_template_kwargs with reasoning_effort set to no_think through extra_body.
Context and sampling
Context is 1M tokens. The card documents no rope or YaRN override. It recommends temperature=0.9 and top_p=1.0.
Pitfalls
- The vLLM example sets a sparse attention backend explicitly.
- Parser names are
hy_v4, not a generic reasoning parser. - The card does not say whether remote code is required, so follow the image's own example command.
Launch commands by Hy4 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Hy4-preview (780.0B)
Needs about 1743 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve tencent/Hy4-previewSGLang
sglang serve --model-path tencent/Hy4-previewGeneric example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/tencent/Hy4-preview
- https://huggingface.co/tencent/Hy4-preview/blob/main/README.md
Updated 2026-10-07.