How to run DeepSeek-V4 with vLLM and SGLang

Engine support, custom chat encoding, the three thinking modes and sampling settings for running DeepSeek-V4 Pro and Flash.

This guide covers serving the DeepSeek-V4 family. Sizing and launch commands per model appear below.

Engine support

The DeepSeek-V4-Pro card lists Transformers, vLLM and SGLang, each with an OpenAI-compatible API. The card page I read states no minimum engine version, and it does not mention llama.cpp or Ollama. Check the card for current guidance.

Models and precision

  • DeepSeek-V4-Pro: 1.6T total parameters, 49B activated.
  • DeepSeek-V4-Flash: 284B total parameters, 13B activated.
  • Both support a one million token context.
  • Weights are mixed precision: MoE expert parameters use FP4, most other parameters use FP8 (FP8).

Chat encoding

The model does not use a Jinja chat template. The card points to its encoding folder for converting OpenAI-compatible messages into model input and for parsing the output. Use that code rather than a generic template call.

Thinking modes

Three reasoning effort levels are supported:

  • Non-think: fast responses for routine tasks.
  • Think High: slower, more accurate analysis for complex problems.
  • Think Max: pushes reasoning as far as it goes. The card says it needs a context window of at least 384K tokens.

Sampling

For local deployment the card recommends temperature = 1.0 and top_p = 1.0.

Pitfalls

  • Set the maximum context high enough before using Think Max.
  • Sending messages through a generic template instead of the encoding code will not match the format the model expects.

Launch commands by DeepSeek V4 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

DeepSeek-V4-Flash-0731 (304.2B)

Needs about 564 GB of VRAM at FP4 + FP8 Mixed. Cheapest live fit: 6× RTX PRO 6000 at $8.25/hr.

vLLM

vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 --tensor-parallel-size 6

SGLang

sglang serve --model-path deepseek-ai/DeepSeek-V4-Flash-0731 --tp 6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

DeepSeek-V4-Pro-0813 (1650.5B)

Needs about 1128 GB of VRAM at FP4 + FP8 Mixed. No live GPU fit right now.

vLLM

vllm serve deepseek-ai/DeepSeek-V4-Pro-0813

SGLang

sglang serve --model-path deepseek-ai/DeepSeek-V4-Pro-0813

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.