How to run Hy4 with vLLM and SGLang

Serving Tencent Hy4-preview with vLLM and SGLang: prebuilt images, parser flags, MTP speculative decoding and the reasoning_effort toggle.

This guide covers serving the Hy4 family. The checkpoint is named Hy4-preview, with an FP8 build (FP8). Memory sizing and launch commands appear below.

Engine support

The card documents vLLM and SGLang only, each through a prebuilt image: vllm/vllm-openai:hy4-preview and lmsysorg/sglang:hy4-preview (the latter for x86 and ARM). It states no minimum version, so use those images rather than a stock release. It does not cover llama.cpp or Ollama.

vLLM flags

  • --tensor-parallel-size 8
  • --speculative-config with method mtp and 3 speculative tokens
  • --attention-backend FLASHMLA_SPARSE
  • --tool-call-parser hy_v4
  • --reasoning-parser hy_v4

SGLang flags

  • --tp-size 8
  • --reasoning-parser auto
  • --speculative-algorithm NEXTN with --speculative-num-steps 3

Reasoning toggle

Reasoning defaults to high. For direct answers pass chat_template_kwargs with reasoning_effort set to no_think through extra_body.

Context and sampling

Context is 1M tokens. The card documents no rope or YaRN override. It recommends temperature=0.9 and top_p=1.0.

Pitfalls

  • The vLLM example sets a sparse attention backend explicitly.
  • Parser names are hy_v4, not a generic reasoning parser.
  • The card does not say whether remote code is required, so follow the image's own example command.

Launch commands by Hy4 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Hy4-preview (780.0B)

Needs about 1743 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve tencent/Hy4-preview

SGLang

sglang serve --model-path tencent/Hy4-preview

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.