How to run Qwen3.5 with vLLM, SGLang and llama.cpp
Engine support, thinking toggle, YaRN long context and parser flags for running Qwen3.5 with vLLM, SGLang and Transformers.
This guide covers serving the Qwen3.5 family. Sizing and launch commands per model appear below.
Engine support
The Qwen3.5-35B-A3B card lists SGLang, vLLM, KTransformers and Hugging Face Transformers. At release time it pointed to the main branch of SGLang and Transformers and a nightly vLLM build, with no fixed minimum release stated. Check the card for current guidance. The card does not mention llama.cpp or Ollama.
Thinking toggle and parsers
- Thinking is on by default, with output inside
<think>tags. - Turn it off with
"chat_template_kwargs": {"enable_thinking": False}. - Use
--reasoning-parser qwen3and--tool-call-parser qwen3_coder. --speculative-algo NEXTNenables multi-token prediction.
Long context
Native context is 262,144 tokens, extendable to about 1,010,000 with YaRN (rope scaling). In config.json rope_parameters the card shows rope_type "yarn", factor 4.0 and original_max_position_embeddings 262144. Static YaRN can hurt short-text quality, so set it only when needed.
The card advises keeping context at 128K tokens or more to preserve thinking quality, and warns that the default maximum context can run out of memory.
Multimodal
The model is a unified vision-language model that accepts images and video. Video frame limits can be changed with a configuration override described on the card.
Pitfalls
Launch commands by Qwen3.5 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen-AgentWorld-35B-A3B (34.7B)
Needs about 77.5 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve Qwen/Qwen-AgentWorld-35B-A3B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen-AgentWorld-35B-A3B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/Qwen/Qwen3.5-35B-A3B
- https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html
Updated 2026-10-07.