How to run DeepSeek-R1 with vLLM and SGLang

Prompting rules, temperature, think-tag handling and engine commands for DeepSeek-R1 and its distilled models.

This guide covers serving the DeepSeek-R1 family. Sizing and launch commands per model appear below.

Engine support

The model card gives example commands for vLLM and SGLang, using a distilled model (DeepSeek-R1-Distill-Qwen-32B). The SGLang example passes --trust-remote-code. The card states no minimum engine versions and does not mention llama.cpp or Ollama.

Context length

DeepSeek-R1 and DeepSeek-R1-Zero support a 128K context window. The vLLM example on the card caps --max-model-len at 32768 for the distilled model, so raise it deliberately if you need more.

Prompting recommendations

The card lists these usage recommendations:

  • Set temperature in the 0.5 to 0.7 range (0.6 recommended) to avoid repetition or incoherent output.
  • Do not add a system prompt. Put all instructions in the user prompt.
  • For math, add a directive such as "Please reason step by step, and put your final answer within \boxed."
  • Force the response to begin with <think>\n so the model reasons thoroughly.

Distilled models

There are six distilled models: Qwen-based at 1.5B, 7B, 14B and 32B, and Llama-based at 8B and 70B. They are dense models that serve like their base architectures, so engine support follows the Qwen2.5 and Llama checkpoints they came from.

Pitfalls

  • A system prompt can degrade results, which trips up OpenAI-style clients that always send one.
  • Output contains the reasoning before the answer, so parse or strip the <think> section in your application.

Launch commands by DeepSeek R1 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

DeepSeek-R1-0528-Qwen3-8B (8.2B)

Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve deepseek-ai/DeepSeek-R1-0528-Qwen3-8B --tensor-parallel-size 1

SGLang

sglang serve --model-path deepseek-ai/DeepSeek-R1-0528-Qwen3-8B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

DeepSeek-R1 (684.5B)

Needs about 765 GB of VRAM at F8_E4M3. Cheapest live fit: 8× RTX PRO 6000 at $11.70/hr.

vLLM

vllm serve deepseek-ai/DeepSeek-R1 --tensor-parallel-size 8

SGLang

sglang serve --model-path deepseek-ai/DeepSeek-R1 --tp 8

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.