How to run Qwen1.5 with vLLM, SGLang and llama.cpp

Transformers requirement, chat template, 32K context and no trust-remote-code notes for running Qwen1.5 chat models.

This guide covers serving the Qwen1.5 family. Exact memory sizing and launch commands for each size appear below.

Library requirements

  • The Qwen1.5-7B-Chat card advises transformers>=4.37.0. Older versions raise KeyError: 'qwen2'.
  • The card says trust_remote_code is not needed, unlike earlier Qwen releases.
  • The card does not state minimum versions for vLLM, SGLang, llama.cpp or Ollama, so check each engine's own release notes.

Chat template

The chat checkpoints include a chat template. Apply it with tokenizer.apply_chat_template() before generating, or let your server do so on chat endpoints.

Context length

The card states 32K context length for models of all sizes. It does not describe a YaRN or rope-scaling setup for going further, so keep prompts within 32K.

Pitfalls

  • If you see code switching (the model drifting between languages) or other bad cases, the card advises using the hyper-parameters in the checkpoint's generation_config.json.
  • Do not pass --trust-remote-code out of habit from older Qwen guides: it is not required here.

Launch commands by Qwen1.5 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Qwen1.5-0.5B-Chat (620M)

Needs about 1.4 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen1.5-0.5B-Chat --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-0.5B-Chat --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-1.8B-Chat (1.8B)

Needs about 4.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen1.5-1.8B-Chat --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-1.8B-Chat --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-4B (4.0B)

Needs about 8.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve Qwen/Qwen1.5-4B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-4B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

CodeQwen1.5-7B-Chat (7.3B)

Needs about 16.2 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve Qwen/CodeQwen1.5-7B-Chat --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/CodeQwen1.5-7B-Chat --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-7B (7.7B)

Needs about 17.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve Qwen/Qwen1.5-7B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-7B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-14B (14.2B)

Needs about 31.7 GB of VRAM at BF16. Cheapest live fit: RTX 4080 Super at $0.338/hr.

vLLM

vllm serve Qwen/Qwen1.5-14B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-14B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-MoE-A2.7B (14.3B)

Needs about 32.0 GB of VRAM at BF16. Cheapest live fit: RTX 4080 Super at $0.338/hr.

vLLM

vllm serve Qwen/Qwen1.5-MoE-A2.7B --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-MoE-A2.7B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-32B-Chat (32.5B)

Needs about 72.7 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve Qwen/Qwen1.5-32B-Chat --tensor-parallel-size 1

SGLang

sglang serve --model-path Qwen/Qwen1.5-32B-Chat --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Qwen1.5-72B-Chat (72.3B)

Needs about 162 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve Qwen/Qwen1.5-72B-Chat --tensor-parallel-size 7

SGLang

sglang serve --model-path Qwen/Qwen1.5-72B-Chat --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.