How to run Qwen2.5 with vLLM, SGLang and llama.cpp

Transformers version, YaRN long-context config and chat template notes for running Qwen2.5 with vLLM and other engines.

This guide covers serving the Qwen2.5 family. Exact memory sizing and launch commands for each size appear below.

Engine and library requirements

  • The model cards require transformers>=4.37.0. Older versions fail with KeyError: 'qwen2', and the cards recommend the latest release.
  • The Qwen2.5-7B-Instruct card points to vLLM for long-context deployment and links the Qwen vLLM deployment docs.
  • The cards I opened do not state minimum versions for SGLang, llama.cpp or Ollama, so check each engine's own release notes.

Chat template

Instruct checkpoints ship a chat template with system, user and assistant roles. Apply it with apply_chat_template(), or let your server apply it on chat requests. The base checkpoints are not meant for chat: the Qwen2.5-7B card says "We do not recommend using base language models for conversations."

Long context

The 7B card lists a 131,072 token context window and up to 8,192 generated tokens, with the default config set for 32,768 tokens. To go past 32K, add YaRN (rope scaling) to config.json:

  • type "yarn"
  • factor 4.0
  • original_max_position_embeddings 32768

Pitfalls

  • vLLM supports only static YaRN: the factor stays constant regardless of input length, which can hurt quality on short texts. Add the setting only when you need long inputs.
  • Keep the generation limit in mind: the card says outputs go up to 8K tokens.

Launch commands by Qwen2.5 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen2.5-0.5B-Instruct (494M)

Needs about 1.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen2.5-0.5B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-0.5B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2.5-1.5B-Instruct (1.5B)

Needs about 3.5 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen2.5-1.5B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-1.5B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2.5-3B-Instruct (3.1B)

Needs about 6.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen2.5-3B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-3B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen2.5:3b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf bartowski/Qwen2.5-3B-Instruct-GGUF

Prebuilt GGUF weights published at bartowski/Qwen2.5-3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen2.5-7B-Instruct (7.6B)

Needs about 17.0 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve Qwen/Qwen2.5-7B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-7B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen2.5:7b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf bartowski/Qwen2.5-7B-Instruct-GGUF

Prebuilt GGUF weights published at bartowski/Qwen2.5-7B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen2.5-14B-Instruct (14.8B)

Needs about 33.0 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.

vLLM

vllm serve Qwen/Qwen2.5-14B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-14B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run qwen2.5:14b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf bartowski/Qwen2.5-14B-Instruct-GGUF

Prebuilt GGUF weights published at bartowski/Qwen2.5-14B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Qwen2.5-32B-Instruct (32.8B)

Needs about 73.2 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve Qwen/Qwen2.5-32B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen2.5-32B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen2.5-72B-Instruct (72.7B)

Needs about 163 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 7

SGLang

sglang serve --model-path Qwen/Qwen2.5-72B-Instruct --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.