How to run Qwen3.6 with vLLM, SGLang and llama.cpp

Minimum vLLM and SGLang versions, thinking toggle, preserve_thinking, YaRN and parser flags for running Qwen3.6 models.

This guide covers serving the Qwen3.6 family. Sizing and launch commands appear below.

Engine support

The Qwen3.6 model cards state:

  • SGLang 0.5.10 or newer
  • vLLM 0.19.0 or newer
  • Hugging Face Transformers (latest, with serving support)
  • KTransformers (see its deployment guide)

The Qwen3.6-27B page lists quantized builds usable from llama.cpp and Ollama. The cards give no minimum versions for those.

Thinking behavior

  • Thinking is on by default, with <think> blocks before the answer.
  • Disable it with "chat_template_kwargs": {"enable_thinking": False}.
  • Qwen3.6-27B adds preserve_thinking, which keeps reasoning from earlier messages. The card says it helps agent scenarios and can reduce token use.
  • Parser flags: --reasoning-parser qwen3, --tool-call-parser qwen3_coder, and --speculative-algo NEXTN for multi-token prediction.

Long context

Native context is 262,144 tokens, extendable to 1,010,000 with YaRN (rope scaling) by editing rope_parameters in config.json or using command-line overrides. Static YaRN can hurt short-text quality. The card advises keeping context at 128K tokens or more to preserve thinking capability.

Multimodal and pitfalls

  • The model takes text, images and video. For hour-scale video, raise longest_edge to 469,762,048 as described on the card.
  • A presence_penalty above 2.0 may cause language mixing.
  • FP8 variants are published (FP8); see the family hub for sizes.

Launch commands by Qwen3.6 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

No Qwen3.6 model of a size we can compute is in the catalog yet. See the Qwen3.6 model list for what is published.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.