How to run DeepSeek-V4 with vLLM and SGLang
Engine support, custom chat encoding, the three thinking modes and sampling settings for running DeepSeek-V4 Pro and Flash.
This guide covers serving the DeepSeek-V4 family. Sizing and launch commands per model appear below.
Engine support
The DeepSeek-V4-Pro card lists Transformers, vLLM and SGLang, each with an OpenAI-compatible API. The card page I read states no minimum engine version, and it does not mention llama.cpp or Ollama. Check the card for current guidance.
Models and precision
- DeepSeek-V4-Pro: 1.6T total parameters, 49B activated.
- DeepSeek-V4-Flash: 284B total parameters, 13B activated.
- Both support a one million token context.
- Weights are mixed precision: MoE expert parameters use FP4, most other parameters use FP8 (FP8).
Chat encoding
The model does not use a Jinja chat template. The card points to its encoding folder for converting OpenAI-compatible messages into model input and for parsing the output. Use that code rather than a generic template call.
Thinking modes
Three reasoning effort levels are supported:
- Non-think: fast responses for routine tasks.
- Think High: slower, more accurate analysis for complex problems.
- Think Max: pushes reasoning as far as it goes. The card says it needs a context window of at least 384K tokens.
Sampling
For local deployment the card recommends temperature = 1.0 and top_p = 1.0.
Pitfalls
- Set the maximum context high enough before using Think Max.
- Sending messages through a generic template instead of the encoding code will not match the format the model expects.
Launch commands by DeepSeek V4 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
DeepSeek-V4-Flash-0731 (304.2B)
Needs about 564 GB of VRAM at FP4 + FP8 Mixed. Cheapest live fit: 6× RTX PRO 6000 at $8.25/hr.
vLLM
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 --tensor-parallel-size 6SGLang
sglang serve --model-path deepseek-ai/DeepSeek-V4-Flash-0731 --tp 6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
DeepSeek-V4-Pro-0813 (1650.5B)
Needs about 1128 GB of VRAM at FP4 + FP8 Mixed. No live GPU fit right now.
vLLM
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813SGLang
sglang serve --model-path deepseek-ai/DeepSeek-V4-Pro-0813Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.