How to run Qwen1.5 with vLLM, SGLang and llama.cpp
Transformers requirement, chat template, 32K context and no trust-remote-code notes for running Qwen1.5 chat models.
This guide covers serving the Qwen1.5 family. Exact memory sizing and launch commands for each size appear below.
Library requirements
- The Qwen1.5-7B-Chat card advises
transformers>=4.37.0. Older versions raiseKeyError: 'qwen2'. - The card says
trust_remote_codeis not needed, unlike earlier Qwen releases. - The card does not state minimum versions for vLLM, SGLang, llama.cpp or Ollama, so check each engine's own release notes.
Chat template
The chat checkpoints include a chat template. Apply it with tokenizer.apply_chat_template() before generating, or let your server do so on chat endpoints.
Context length
The card states 32K context length for models of all sizes. It does not describe a YaRN or rope-scaling setup for going further, so keep prompts within 32K.
Pitfalls
- If you see code switching (the model drifting between languages) or other bad cases, the card advises using the hyper-parameters in the checkpoint's
generation_config.json. - Do not pass
--trust-remote-codeout of habit from older Qwen guides: it is not required here.
Launch commands by Qwen1.5 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Qwen1.5-0.5B-Chat (620M)
Needs about 1.4 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen1.5-0.5B-Chat --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-0.5B-Chat --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-1.8B-Chat (1.8B)
Needs about 4.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen1.5-1.8B-Chat --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-1.8B-Chat --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-4B (4.0B)
Needs about 8.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve Qwen/Qwen1.5-4B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-4B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
CodeQwen1.5-7B-Chat (7.3B)
Needs about 16.2 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve Qwen/CodeQwen1.5-7B-Chat --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/CodeQwen1.5-7B-Chat --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-7B (7.7B)
Needs about 17.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve Qwen/Qwen1.5-7B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-7B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-14B (14.2B)
Needs about 31.7 GB of VRAM at BF16. Cheapest live fit: RTX 4080 Super at $0.338/hr.
vLLM
vllm serve Qwen/Qwen1.5-14B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-14B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-MoE-A2.7B (14.3B)
Needs about 32.0 GB of VRAM at BF16. Cheapest live fit: RTX 4080 Super at $0.338/hr.
vLLM
vllm serve Qwen/Qwen1.5-MoE-A2.7B --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-MoE-A2.7B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-32B-Chat (32.5B)
Needs about 72.7 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve Qwen/Qwen1.5-32B-Chat --tensor-parallel-size 1SGLang
sglang serve --model-path Qwen/Qwen1.5-32B-Chat --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Qwen1.5-72B-Chat (72.3B)
Needs about 162 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.
vLLM
vllm serve Qwen/Qwen1.5-72B-Chat --tensor-parallel-size 7SGLang
sglang serve --model-path Qwen/Qwen1.5-72B-Chat --tp 7Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.