How to run Qwen3.8 with vLLM, SGLang and llama.cpp

Engine support, thinking and reasoning_effort controls, YaRN long context and pitfalls for running Qwen3.8 models.

This guide covers serving the Qwen3.8 family. Sizing and launch commands appear below.

Engine support

The model cards point to SGLang (cookbook), vLLM (recipe) and TokenSpeed (recipe), plus Hugging Face Transformers. Quantized builds are listed for Ollama, llama.cpp and LM Studio. The cards state no minimum engine versions, so follow the linked recipes and use recent releases.

Thinking controls

  • Qwen3.8-27B and Qwen3.8-Flash-Next think by default. Disable with enable_thinking set to false.
  • preserve_thinking (default true) keeps reasoning from earlier messages.
  • reasoning_effort accepts xhigh (default), medium or low. The cards warn that lower effort in agentic tasks can raise total latency through failed retries.
  • Qwen3.8-2.4T-A95B always thinks and the card says this cannot be disabled.
  • Streamed deltas carry reasoning_content or reasoning fields for separating thoughts from the answer.
  • Sampling for thinking mode on the 27B card: temperature 1.0, top_p 0.95.

Long context

Native context is 262,144 tokens, extendable to roughly 1,000,000 with YaRN (rope scaling). Set rope_type "yarn" and factor 4.0 in rope_parameters. The Flash-Next card also requires VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 for vLLM or SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 for SGLang. Static YaRN may hurt short texts.

Pitfalls

  • Qwen3.8-2.4T-A95B is text-only; multimodal input is not supported. The 27B model accepts images and video.
  • Give agentic tasks adequate output length.
  • A higher presence_penalty may cause language mixing.
  • FP8 builds (FP8) are published for the three main models.

Launch commands by Qwen3.8 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen3.8-2.4T-A95B (2446.2B)

Needs about 5468 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve Qwen/Qwen3.8-2.4T-A95B

SGLang

sglang serve --model-path Qwen/Qwen3.8-2.4T-A95B

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.