How to run Qwen3.5 with vLLM, SGLang and llama.cpp

Engine support, thinking toggle, YaRN long context and parser flags for running Qwen3.5 with vLLM, SGLang and Transformers.

This guide covers serving the Qwen3.5 family. Sizing and launch commands per model appear below.

Engine support

The Qwen3.5-35B-A3B card lists SGLang, vLLM, KTransformers and Hugging Face Transformers. At release time it pointed to the main branch of SGLang and Transformers and a nightly vLLM build, with no fixed minimum release stated. Check the card for current guidance. The card does not mention llama.cpp or Ollama.

Thinking toggle and parsers

  • Thinking is on by default, with output inside <think> tags.
  • Turn it off with "chat_template_kwargs": {"enable_thinking": False}.
  • Use --reasoning-parser qwen3 and --tool-call-parser qwen3_coder.
  • --speculative-algo NEXTN enables multi-token prediction.

Long context

Native context is 262,144 tokens, extendable to about 1,010,000 with YaRN (rope scaling). In config.json rope_parameters the card shows rope_type "yarn", factor 4.0 and original_max_position_embeddings 262144. Static YaRN can hurt short-text quality, so set it only when needed.

The card advises keeping context at 128K tokens or more to preserve thinking quality, and warns that the default maximum context can run out of memory.

Multimodal

The model is a unified vision-language model that accepts images and video. Video frame limits can be changed with a configuration override described on the card.

Pitfalls

  • Release-dependent engine installs: confirm your engine is recent enough.
  • Quantized variants exist, including FP8 (FP8) and GPTQ Int4 (GPTQ) in the Qwen organization listing.

Launch commands by Qwen3.5 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen-AgentWorld-35B-A3B (34.7B)

Needs about 77.5 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve Qwen/Qwen-AgentWorld-35B-A3B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen-AgentWorld-35B-A3B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.