How to run Qwen2.5 with vLLM, SGLang and llama.cpp
Transformers version, YaRN long-context config and chat template notes for running Qwen2.5 with vLLM and other engines.
This guide covers serving the Qwen2.5 family. Exact memory sizing and launch commands for each size appear below.
Engine and library requirements
- The model cards require
transformers>=4.37.0. Older versions fail withKeyError: 'qwen2', and the cards recommend the latest release. - The Qwen2.5-7B-Instruct card points to vLLM for long-context deployment and links the Qwen vLLM deployment docs.
- The cards I opened do not state minimum versions for SGLang, llama.cpp or Ollama, so check each engine's own release notes.
Chat template
Instruct checkpoints ship a chat template with system, user and assistant roles. Apply it with apply_chat_template(), or let your server apply it on chat requests. The base checkpoints are not meant for chat: the Qwen2.5-7B card says "We do not recommend using base language models for conversations."
Long context
The 7B card lists a 131,072 token context window and up to 8,192 generated tokens, with the default config set for 32,768 tokens. To go past 32K, add YaRN (rope scaling) to config.json:
type"yarn"factor4.0original_max_position_embeddings32768
Pitfalls
- vLLM supports only static YaRN: the factor stays constant regardless of input length, which can hurt quality on short texts. Add the setting only when you need long inputs.
- Keep the generation limit in mind: the card says outputs go up to 8K tokens.
Launch commands by Qwen2.5 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen2.5-0.5B-Instruct (494M)
Needs about 1.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen2.5-0.5B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-0.5B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2.5-1.5B-Instruct (1.5B)
Needs about 3.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen2.5-1.5B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-1.5B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2.5-3B-Instruct (3.1B)
Needs about 6.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen2.5-3B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-3B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf bartowski/Qwen2.5-3B-Instruct-GGUFPrebuilt GGUF weights published at bartowski/Qwen2.5-3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen2.5-7B-Instruct (7.6B)
Needs about 17.0 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve Qwen/Qwen2.5-7B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-7B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf bartowski/Qwen2.5-7B-Instruct-GGUFPrebuilt GGUF weights published at bartowski/Qwen2.5-7B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen2.5-14B-Instruct (14.8B)
Needs about 33.0 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.
vLLM
vllm serve Qwen/Qwen2.5-14B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-14B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf bartowski/Qwen2.5-14B-Instruct-GGUFPrebuilt GGUF weights published at bartowski/Qwen2.5-14B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Qwen2.5-32B-Instruct (32.8B)
Needs about 73.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve Qwen/Qwen2.5-32B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen2.5-32B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen2.5-72B-Instruct (72.7B)
Needs about 163 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.
vLLM
vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 7SGLang
sglang serve --model-path Qwen/Qwen2.5-72B-Instruct --tp 7Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.