How to run Qwen3 with vLLM, SGLang and llama.cpp
Engine support, thinking toggle, YaRN long context and parser flags for running Qwen3 with vLLM, SGLang, llama.cpp and Ollama.
This guide covers serving the Qwen3 family. Exact memory sizing and launch commands for each size appear below.
Engine support
The Qwen3-32B model card lists these minimums:
- vLLM 0.8.5 or newer
- SGLang 0.4.6.post1 or newer
- Hugging Face Transformers: use a recent release, since versions below 4.51.0 fail with a
KeyError - llama.cpp and Ollama are also supported
The Qwen docs state that llama.cpp supports Qwen3 and Qwen3MoE from build b5092. Use the GGUF files published by Qwen.
Thinking and non-thinking modes
Qwen3 switches modes inside one chat template.
enable_thinkingdefaults to true. Set it to false for standard chat behavior.- With thinking on, append
/thinkor/no_thinkto a user turn to toggle mid-conversation. - For parsing reasoning out of the reply, the card shows
--reasoning-parser qwen3for SGLang and--enable-reasoning --reasoning-parser deepseek_r1for vLLM. - Do not use greedy decoding in thinking mode: it causes degradation and endless repetition.
- Leave thinking content out of multi-turn history.
Long context
Native context is 32,768 tokens, extendable to 131,072 with YaRN (rope scaling). In config.json set rope_scaling with rope_type "yarn", factor 4.0 and original_max_position_embeddings 32768. Static YaRN can hurt quality on short texts, so enable it only when you need the length.
In llama.cpp the Qwen docs use --rope-scaling yarn --rope-scale 4 --yarn-orig-ctx 32768 with -c 131072. Use --jinja to apply the chat template embedded in the GGUF. llama.cpp does not expose a hard thinking switch, so the docs suggest a custom template file with enable_thinking set to false.
Pitfalls
- Old Transformers versions break loading.
- Reasoning parser flags differ between vLLM and SGLang.
- Adjust the llama.cpp default context (
-cdefaults to 4096).
Launch commands by Qwen3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen3-0.6B (752M)
Needs about 1.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen3-0.6B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-0.6B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen3-1.7B (2.0B)
Needs about 4.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen3-1.7B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-1.7B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Qwen3-1.7B-GGUFPrebuilt GGUF weights published at unsloth/Qwen3-1.7B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen3-4B (4.0B)
Needs about 9.0 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen3-4B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-4B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen3-8B (8.2B)
Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve Qwen/Qwen3-8B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-8B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Qwen3-8B-GGUFPrebuilt GGUF weights published at unsloth/Qwen3-8B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen3-14B (14.8B)
Needs about 33.0 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.
vLLM
vllm serve Qwen/Qwen3-14B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-14B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen3-30B-A3B (30.5B)
Needs about 68.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve Qwen/Qwen3-30B-A3B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-30B-A3B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf KVCache-ai/Qwen3-30BA3B-GGUFPrebuilt GGUF weights published at KVCache-ai/Qwen3-30BA3B-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen3-32B (32.8B)
Needs about 73.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve Qwen/Qwen3-32B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen3-32B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen3-Coder-Next (79.7B)
Needs about 178 GB of VRAM at BF16. Cheapest live fit: 8× RTX A5000 at $1.41/hr.
vLLM
vllm serve Qwen/Qwen3-Coder-Next --tensor-parallel-size 8SGLang
sglang serve --model-path Qwen/Qwen3-Coder-Next --tp 8Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Qwen3-Coder-Next-GGUFPrebuilt GGUF weights published at unsloth/Qwen3-Coder-Next-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen3-Next-80B-A3B-Instruct (81.3B)
Needs about 182 GB of VRAM at BF16. Cheapest live fit: 8× RTX A5000 at $1.41/hr.
vLLM
vllm serve Qwen/Qwen3-Next-80B-A3B-Instruct --tensor-parallel-size 8SGLang
sglang serve --model-path Qwen/Qwen3-Next-80B-A3B-Instruct --tp 8Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Qwen3-Next-80B-A3B-Instruct-GGUFPrebuilt GGUF weights published at unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen3-235B-A22B (235.1B)
Needs about 525 GB of VRAM at BF16. Cheapest live fit: 6× RTX PRO 6000 at $8.25/hr.
vLLM
vllm serve Qwen/Qwen3-235B-A22B --tensor-parallel-size 6SGLang
sglang serve --model-path Qwen/Qwen3-235B-A22B --tp 6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen3-Coder-480B-A35B-Instruct (480.2B)
Needs about 1073 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve Qwen/Qwen3-Coder-480B-A35B-InstructSGLang
sglang serve --model-path Qwen/Qwen3-Coder-480B-A35B-InstructGeneric example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUFPrebuilt GGUF weights published at unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/Qwen/Qwen3-32B
- https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html
Updated 2026-10-07.