How to run Qwen2 with vLLM, SGLang and llama.cpp
Minimum vLLM and Transformers versions, YaRN long-context config and chat template notes for running Qwen2 locally.
This guide covers serving the Qwen2 family. Exact memory sizing and launch commands for each size appear below.
Engine and library requirements
The Qwen2-7B-Instruct card states:
transformers>=4.37.0. Without it you getKeyError: 'qwen2'.vllm>=0.4.3for vLLM.
The card does not state minimum versions for SGLang, llama.cpp or Ollama, so check each engine's own release notes.
Chat template
The model formats conversations with apply_chat_template, using system, user and assistant roles. Chat servers apply the same template on chat endpoints.
Long context
The card lists a context length of up to 131,072 tokens. For long inputs, add a rope_scaling block to config.json with type "yarn" and factor 4.0 (see rope scaling).
Pitfalls
- vLLM supports only static YaRN: the scaling factor is constant regardless of input length, which can hurt performance on shorter texts. The card says to add
rope_scalingonly when you need long contexts. - Old Transformers versions fail at load time with the
qwen2KeyError.
Launch commands by Qwen2 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen2-0.5B (494M)
Needs about 1.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen2-0.5B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2-0.5B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2-1.5B-Instruct (1.5B)
Needs about 3.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen2-1.5B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2-1.5B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2-7B-Instruct (7.6B)
Needs about 17.0 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve Qwen/Qwen2-7B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2-7B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2-57B-A14B-Instruct (57.4B)
Needs about 128 GB of VRAM at BF16. Cheapest live fit: 6× RTX A5000 at $1.06/hr.
vLLM
vllm serve Qwen/Qwen2-57B-A14B-Instruct --tensor-parallel-size 6SGLang
sglang serve --model-path Qwen/Qwen2-57B-A14B-Instruct --tp 6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2-72B-Instruct (72.7B)
Needs about 163 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.
vLLM
vllm serve Qwen/Qwen2-72B-Instruct --tensor-parallel-size 7SGLang
sglang serve --model-path Qwen/Qwen2-72B-Instruct --tp 7Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.