How to run Qwen3 with vLLM, SGLang and llama.cpp

Engine support, thinking toggle, YaRN long context and parser flags for running Qwen3 with vLLM, SGLang, llama.cpp and Ollama.

This guide covers serving the Qwen3 family. Exact memory sizing and launch commands for each size appear below.

Engine support

The Qwen3-32B model card lists these minimums:

  • vLLM 0.8.5 or newer
  • SGLang 0.4.6.post1 or newer
  • Hugging Face Transformers: use a recent release, since versions below 4.51.0 fail with a KeyError
  • llama.cpp and Ollama are also supported

The Qwen docs state that llama.cpp supports Qwen3 and Qwen3MoE from build b5092. Use the GGUF files published by Qwen.

Thinking and non-thinking modes

Qwen3 switches modes inside one chat template.

  • enable_thinking defaults to true. Set it to false for standard chat behavior.
  • With thinking on, append /think or /no_think to a user turn to toggle mid-conversation.
  • For parsing reasoning out of the reply, the card shows --reasoning-parser qwen3 for SGLang and --enable-reasoning --reasoning-parser deepseek_r1 for vLLM.
  • Do not use greedy decoding in thinking mode: it causes degradation and endless repetition.
  • Leave thinking content out of multi-turn history.

Long context

Native context is 32,768 tokens, extendable to 131,072 with YaRN (rope scaling). In config.json set rope_scaling with rope_type "yarn", factor 4.0 and original_max_position_embeddings 32768. Static YaRN can hurt quality on short texts, so enable it only when you need the length.

In llama.cpp the Qwen docs use --rope-scaling yarn --rope-scale 4 --yarn-orig-ctx 32768 with -c 131072. Use --jinja to apply the chat template embedded in the GGUF. llama.cpp does not expose a hard thinking switch, so the docs suggest a custom template file with enable_thinking set to false.

Pitfalls

  • Old Transformers versions break loading.
  • Reasoning parser flags differ between vLLM and SGLang.
  • Adjust the llama.cpp default context (-c defaults to 4096).

Launch commands by Qwen3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen3-0.6B (752M)

Needs about 1.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen3-0.6B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-0.6B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen3-1.7B (2.0B)

Needs about 4.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen3-1.7B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-1.7B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3:1.7b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Qwen3-1.7B-GGUF

Prebuilt GGUF weights published at unsloth/Qwen3-1.7B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen3-4B (4.0B)

Needs about 9.0 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen3-4B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-4B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3:4b

Verified against Ollama's own library listing. Source

Qwen3-8B (8.2B)

Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve Qwen/Qwen3-8B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-8B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3:8b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Qwen3-8B-GGUF

Prebuilt GGUF weights published at unsloth/Qwen3-8B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen3-14B (14.8B)

Needs about 33.0 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.

vLLM

vllm serve Qwen/Qwen3-14B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-14B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen3-30B-A3B (30.5B)

Needs about 68.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve Qwen/Qwen3-30B-A3B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-30B-A3B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3:30b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf KVCache-ai/Qwen3-30BA3B-GGUF

Prebuilt GGUF weights published at KVCache-ai/Qwen3-30BA3B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen3-32B (32.8B)

Needs about 73.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve Qwen/Qwen3-32B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen3-32B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3:32b

Verified against Ollama's own library listing. Source

Qwen3-Coder-Next (79.7B)

Needs about 178 GB of VRAM at BF16. Cheapest live fit: 8× RTX A5000 at $1.41/hr.

vLLM

vllm serve Qwen/Qwen3-Coder-Next --tensor-parallel-size 8

SGLang

sglang serve --model-path Qwen/Qwen3-Coder-Next --tp 8

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3-coder-next

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF

Prebuilt GGUF weights published at unsloth/Qwen3-Coder-Next-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen3-Next-80B-A3B-Instruct (81.3B)

Needs about 182 GB of VRAM at BF16. Cheapest live fit: 8× RTX A5000 at $1.41/hr.

vLLM

vllm serve Qwen/Qwen3-Next-80B-A3B-Instruct --tensor-parallel-size 8

SGLang

sglang serve --model-path Qwen/Qwen3-Next-80B-A3B-Instruct --tp 8

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3-next:80b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF

Prebuilt GGUF weights published at unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen3-235B-A22B (235.1B)

Needs about 525 GB of VRAM at BF16. Cheapest live fit: 6× RTX PRO 6000 at $8.25/hr.

vLLM

vllm serve Qwen/Qwen3-235B-A22B --tensor-parallel-size 6

SGLang

sglang serve --model-path Qwen/Qwen3-235B-A22B --tp 6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen3-Coder-480B-A35B-Instruct (480.2B)

Needs about 1073 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve Qwen/Qwen3-Coder-480B-A35B-Instruct

SGLang

sglang serve --model-path Qwen/Qwen3-Coder-480B-A35B-Instruct

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen3-coder:480b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF

Prebuilt GGUF weights published at unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.