How to run Qwen2 with vLLM, SGLang and llama.cpp

Minimum vLLM and Transformers versions, YaRN long-context config and chat template notes for running Qwen2 locally.

This guide covers serving the Qwen2 family. Exact memory sizing and launch commands for each size appear below.

Engine and library requirements

The Qwen2-7B-Instruct card states:

  • transformers>=4.37.0. Without it you get KeyError: 'qwen2'.
  • vllm>=0.4.3 for vLLM.

The card does not state minimum versions for SGLang, llama.cpp or Ollama, so check each engine's own release notes.

Chat template

The model formats conversations with apply_chat_template, using system, user and assistant roles. Chat servers apply the same template on chat endpoints.

Long context

The card lists a context length of up to 131,072 tokens. For long inputs, add a rope_scaling block to config.json with type "yarn" and factor 4.0 (see rope scaling).

Pitfalls

  • vLLM supports only static YaRN: the scaling factor is constant regardless of input length, which can hurt performance on shorter texts. The card says to add rope_scaling only when you need long contexts.
  • Old Transformers versions fail at load time with the qwen2 KeyError.

Launch commands by Qwen2 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen2-0.5B (494M)

Needs about 1.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen2-0.5B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2-0.5B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2-1.5B-Instruct (1.5B)

Needs about 3.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen2-1.5B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2-1.5B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2-7B-Instruct (7.6B)

Needs about 17.0 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve Qwen/Qwen2-7B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2-7B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2-57B-A14B-Instruct (57.4B)

Needs about 128 GB of VRAM at BF16. Cheapest live fit: 6× RTX A5000 at $1.06/hr.

vLLM

vllm serve Qwen/Qwen2-57B-A14B-Instruct --tensor-parallel-size 6

SGLang

sglang serve --model-path Qwen/Qwen2-57B-A14B-Instruct --tp 6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2-72B-Instruct (72.7B)

Needs about 163 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve Qwen/Qwen2-72B-Instruct --tensor-parallel-size 7

SGLang

sglang serve --model-path Qwen/Qwen2-72B-Instruct --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.