How to run Qwen3.6 with vLLM, SGLang and llama.cpp
Minimum vLLM and SGLang versions, thinking toggle, preserve_thinking, YaRN and parser flags for running Qwen3.6 models.
This guide covers serving the Qwen3.6 family. Sizing and launch commands appear below.
Engine support
The Qwen3.6 model cards state:
- SGLang 0.5.10 or newer
- vLLM 0.19.0 or newer
- Hugging Face Transformers (latest, with serving support)
- KTransformers (see its deployment guide)
The Qwen3.6-27B page lists quantized builds usable from llama.cpp and Ollama. The cards give no minimum versions for those.
Thinking behavior
- Thinking is on by default, with
<think>blocks before the answer. - Disable it with
"chat_template_kwargs": {"enable_thinking": False}. - Qwen3.6-27B adds
preserve_thinking, which keeps reasoning from earlier messages. The card says it helps agent scenarios and can reduce token use. - Parser flags:
--reasoning-parser qwen3,--tool-call-parser qwen3_coder, and--speculative-algo NEXTNfor multi-token prediction.
Long context
Native context is 262,144 tokens, extendable to 1,010,000 with YaRN (rope scaling) by editing rope_parameters in config.json or using command-line overrides. Static YaRN can hurt short-text quality. The card advises keeping context at 128K tokens or more to preserve thinking capability.
Multimodal and pitfalls
- The model takes text, images and video. For hour-scale video, raise
longest_edgeto 469,762,048 as described on the card. - A
presence_penaltyabove 2.0 may cause language mixing. - FP8 variants are published (FP8); see the family hub for sizes.
Launch commands by Qwen3.6 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
No Qwen3.6 model of a size we can compute is in the catalog yet. See the Qwen3.6 model list for what is published.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.