How to run Qwen3.8 with vLLM, SGLang and llama.cpp
Engine support, thinking and reasoning_effort controls, YaRN long context and pitfalls for running Qwen3.8 models.
This guide covers serving the Qwen3.8 family. Sizing and launch commands appear below.
Engine support
The model cards point to SGLang (cookbook), vLLM (recipe) and TokenSpeed (recipe), plus Hugging Face Transformers. Quantized builds are listed for Ollama, llama.cpp and LM Studio. The cards state no minimum engine versions, so follow the linked recipes and use recent releases.
Thinking controls
- Qwen3.8-27B and Qwen3.8-Flash-Next think by default. Disable with
enable_thinkingset to false. preserve_thinking(default true) keeps reasoning from earlier messages.reasoning_effortaccepts xhigh (default), medium or low. The cards warn that lower effort in agentic tasks can raise total latency through failed retries.- Qwen3.8-2.4T-A95B always thinks and the card says this cannot be disabled.
- Streamed deltas carry
reasoning_contentorreasoningfields for separating thoughts from the answer. - Sampling for thinking mode on the 27B card: temperature 1.0, top_p 0.95.
Long context
Native context is 262,144 tokens, extendable to roughly 1,000,000 with YaRN (rope scaling). Set rope_type "yarn" and factor 4.0 in rope_parameters. The Flash-Next card also requires VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 for vLLM or SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 for SGLang. Static YaRN may hurt short texts.
Pitfalls
- Qwen3.8-2.4T-A95B is text-only; multimodal input is not supported. The 27B model accepts images and video.
- Give agentic tasks adequate output length.
- A higher
presence_penaltymay cause language mixing. - FP8 builds (FP8) are published for the three main models.
Launch commands by Qwen3.8 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen3.8-2.4T-A95B (2446.2B)
Needs about 5468 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve Qwen/Qwen3.8-2.4T-A95BSGLang
sglang serve --model-path Qwen/Qwen3.8-2.4T-A95BGeneric example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/Qwen/Qwen3.8-27B
- https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
Updated 2026-10-07.