What GPU do I need to run deepseek-ai/DeepSeek-V4-Pro-0813?
A 1.6T-parameter (49B active) reasoning model with a 1M-token context. 1650.5B parameters, published in FP4 + FP8 Mixed. View on Hugging Face
What DeepSeek-V4-Pro-0813 is
DeepSeek-V4-Pro-0813 is DeepSeek's official release of DeepSeek-V4-Pro, a 1.6T-parameter mixture-of-experts reasoning model with 49B active parameters and a DSpark speculative-decoding module attached, superseding the preview version with improved agentic performance. Its own model card states a 1,048,576-token (1M) context window via a hybrid Compressed/Heavily-Compressed attention architecture, three reasoning-effort levels (low, high, max), and mixed FP4 (MoE experts) + FP8 (other parameters) published precision.
License note: fully permissive, no gating and no usage restrictions. Facts in this section are sourced from DeepSeek-V4-Pro-0813's Hugging Face model card, not benchmarked by Aquanode.
What it's used for
- Multi-step reasoning and agentic coding
- Long-context analysis (up to 1M tokens)
- Tool use and terminal/agent tasks
Benchmarks (published by DeepSeek)
Published by DeepSeek, not measured by Aquanode.
VRAM required & cheapest live GPU fit
Required VRAM = weight size at each precision, plus a fixed overhead for KV-cache, activations, and fragmentation. Full formula and assumptions: methodology.
| Precision | Weight size | Required VRAM | Cheapest live fit | GPUs needed | Est. $/hr (full fit) |
|---|---|---|---|---|---|
| FP4 + FP8 Mixed (native) | 831.4 GB | 1128.0 GB | No capable live offer found | – | – |
A GPU is only matched to a row if its hardware supports that precision, and the primary recommendation is always a single-GPU fit when one exists.
The SGLang DeepSeek-V4 cookbook linked from DeepSeek's model card lists DeepSeek-V4-Pro-0813 as verified on 8×H200 (FP4, tensor-parallel-8); NVIDIA's own H200 spec lists 141GB HBM3e per GPU, so 8×141GB ≈ 1128GB total.
How to run DeepSeek-V4-Pro-0813
Run DeepSeek-V4-Pro-0813 with vLLM
From DeepSeek-V4-Pro-0813's own model card: an example serving it with vLLM and DSpark speculative decoding on a single 4×GB300 node.
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \
--trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
--data-parallel-size 4 --enable-expert-parallel \
--moe-backend deep_gemm_mega_moe \
--attention-config '{"use_fp4_indexer_cache": true}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813/raw/main/README.md
Run DeepSeek-V4-Pro-0813 with SGLang
From DeepSeek-V4-Pro-0813's own model card, same single 4×GB300 node example as the vLLM command above.
sglang serve \
--trust-remote-code \
--model-path deepseek-ai/DeepSeek-V4-Pro-0813 \
--tp 4 \
--moe-runner-backend flashinfer_mxfp4 \
--speculative-algorithm DSPARK \
--mem-fraction-static 0.90 \
--chunked-prefill-size 4096Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813/raw/main/README.md
Deploy DeepSeek-V4-Pro-0813 on Aquanode
Aquanode has no one-click deploy template for DeepSeek-V4-Pro-0813; you install the inference engine yourself with the commands below. Aquanode sells GPU pods billed per second, not a hosted inference API.
- Launch a bare GPU pod sized to the requirement above (1128 GB VRAM or more).
- Open a terminal on the pod, or save one of the commands above as a startup script so it runs automatically the first time the pod boots.
- Run the command and connect to the resulting endpoint.
Weight-to-VRAM math, the fit rules, and how live prices are normalized: full methodology.
More DeepSeek V4 models
- DeepSeek-V4-Flash-0731 (304.2B, FP4 + FP8 Mixed)
Related reading: H100 pricing and specs, The best GPUs for AI, ranked, and Best GPU for LLM inference.