How to run DeepSeek-R1 with vLLM and SGLang
Prompting rules, temperature, think-tag handling and engine commands for DeepSeek-R1 and its distilled models.
This guide covers serving the DeepSeek-R1 family. Sizing and launch commands per model appear below.
Engine support
The model card gives example commands for vLLM and SGLang, using a distilled model (DeepSeek-R1-Distill-Qwen-32B). The SGLang example passes --trust-remote-code. The card states no minimum engine versions and does not mention llama.cpp or Ollama.
Context length
DeepSeek-R1 and DeepSeek-R1-Zero support a 128K context window. The vLLM example on the card caps --max-model-len at 32768 for the distilled model, so raise it deliberately if you need more.
Prompting recommendations
The card lists these usage recommendations:
- Set temperature in the 0.5 to 0.7 range (0.6 recommended) to avoid repetition or incoherent output.
- Do not add a system prompt. Put all instructions in the user prompt.
- For math, add a directive such as "Please reason step by step, and put your final answer within \boxed."
- Force the response to begin with
<think>\nso the model reasons thoroughly.
Distilled models
There are six distilled models: Qwen-based at 1.5B, 7B, 14B and 32B, and Llama-based at 8B and 70B. They are dense models that serve like their base architectures, so engine support follows the Qwen2.5 and Llama checkpoints they came from.
Pitfalls
- A system prompt can degrade results, which trips up OpenAI-style clients that always send one.
- Output contains the reasoning before the answer, so parse or strip the
<think>section in your application.
Launch commands by DeepSeek R1 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
DeepSeek-R1-0528-Qwen3-8B (8.2B)
Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve deepseek-ai/DeepSeek-R1-0528-Qwen3-8B --tensor-parallel-size 1SGLang
sglang serve --model-path deepseek-ai/DeepSeek-R1-0528-Qwen3-8B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
DeepSeek-R1 (684.5B)
Needs about 765 GB of VRAM at F8_E4M3. Cheapest live fit: 8× RTX PRO 6000 at $11.70/hr.
vLLM
vllm serve deepseek-ai/DeepSeek-R1 --tensor-parallel-size 8SGLang
sglang serve --model-path deepseek-ai/DeepSeek-R1 --tp 8Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/deepseek-ai/DeepSeek-R1
- https://huggingface.co/deepseek-ai/DeepSeek-R1/raw/main/README.md
Updated 2026-10-07.