DeepSeek model series
DeepSeek's V2, V3 and V4 mixture-of-experts line plus R1 and side lines: release months, sizes, context length, licenses, and which to use today.
What DeepSeek is
DeepSeek is a Chinese AI lab that publishes open weights under the deepseek-ai organization on Hugging Face. Its main line is large mixture-of-experts (MoE) language models, where only a small share of parameters runs for each token. That keeps serving cost closer to a mid-size model, but every expert still has to sit in GPU memory, so these are multi-GPU deployments. Where a card does not state a date, release months below come from Hugging Face repository creation dates, and V2 is dated from its model card.
Generations in order
DeepSeek-V2 (May 2024). A 236B total, 21B active MoE with Multi-head Latent Attention (MLA) and the DeepSeekMoE design, 128K token context. The card says it cut KV cache by 93.3% against DeepSeek 67B. License: the DeepSeek Model License, which permits commercial use.
DeepSeek-V3 (December 2024). 671B total, 37B active, MLA plus DeepSeekMoE, 128K token context, trained on 14.8 trillion tokens with FP8 mixed precision. The card gives a custom model license that supports commercial use. The V3.1 card lists the same 671B and 37B figures under the MIT license.
DeepSeek-V3.1 and V3.2 (August and December 2025). These sit in the DeepSeek-V3 hub alongside the March 2025 V3-0324 update, V3.1-Terminus (September 2025) and the V3.2-Exp preview (September 2025). V3.1 (August 2025) adds a hybrid thinking mode switched through the chat template, with a 128K context and a longer long-context training phase. V3.2 (December 2025, 685B parameters on the card) introduces DeepSeek Sparse Attention to cut cost at long sequence lengths, and is MIT licensed.
DeepSeek-V4 (April 2026). Two sizes, both with a 1 million token context, MIT license and a hybrid compressed attention design (V4-Pro uses mixed FP4 and FP8 precision):
- V4-Pro: 1.6 trillion total, 49B active, pre-trained on over 32 trillion tokens.
- V4-Flash: 284B total, 13B active.
The V4 cards credit a hybrid of compressed and heavily compressed attention, manifold-constrained hyper-connections and the Muon optimizer for efficiency at long context, and V4-Flash's card says V4-Pro needs 27% of the single-token compute and 10% of the KV cache of its predecessor at a million tokens. Both offer three reasoning levels: non-think, think high and think max. Later updates followed: V4-Flash-0731 (July 2026), V4-Pro-0813 (August 2026) and V4.1-Flash (September 2026). The V4.1-Flash card lists 763B total parameters including 196B of Engram memory, 8B active during prefill and 16B during decode, a vision encoder, a 1 million token context and the MIT license.
Which generation to use today
The organization's newest uploads are the V4 line, so that is the current generation. Pick by license, total parameter count and context length, since those are the facts the cards give.
- Frontier-scale open weights: V4-Pro, if you have a multi-node budget. The card describes it as the best open-source model for reasoning, which is DeepSeek's own claim.
- Smaller V4 deployment: V4-Flash has 13B active parameters and 284B total, and the V4.1-Flash card reports a KV cache about 4x smaller than V4-Flash.
- Permissive license: V3.1 onward and R1 are MIT licensed. V2 and V3 use DeepSeek's own license, so check them before reuse.
- Long context: V4 is the only main-line generation with a 1 million token context; V2, V3 and V3.1 stop at 128K.
- One or two GPUs: none of the main line fits. Use the R1 distilled models instead, in sizes from 1.5B to 70B, built on Qwen2.5 and Llama 3 bases (January 2025, MIT).
Side lines
- DeepSeek-R1: the January 2025 reasoning model. 671B total, 37B active, 128K context, MIT, built on DeepSeek-V3-Base, with six distilled dense models.
- DeepSeek-Coder: the code-specialized models, from the 1.3B and 6.7B deepseek-coder checkpoints to DeepSeek-Coder-V2-Lite-Instruct.
- DeepSeek-MoE: the 16B mixture-of-experts base and chat models, the smallest MoE checkpoints in the organization.
- DeepSeek-VL: vision-language models, including deepseek-vl2 (December 2024).
- DeepSeek-OCR: a 3B vision-language model for OCR and document understanding, MIT licensed.
DeepSeek models can run on Aquanode GPUs.
DeepSeek generations
Every DeepSeek generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| DeepSeek V2 | 5 | 15.7B to 235.7B | 35.1 GB | RTX A6000 | $0.363/hr |
| DeepSeek V3 | 7 | 684.5B to 685.4B | 765 GB | RTX PRO 6000 × 8 | $11.75/hr |
| DeepSeek V4 | 2 | 304.2B to 1650.5B | 564 GB | RTX PRO 6000 × 6 | $8.25/hr |
Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.
Other DeepSeek lines
Side lines from DeepSeek: specialised or companion models outside the main DeepSeek generations.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| DeepSeek R1 | 13 | 1.8B to 684.5B | 18.3 GB | RTX A5000 | $0.176/hr |
| DeepSeek Coder | 6 | 1.3B to 15.7B | 3.0 GB | RTX 4070 Super | $0.121/hr |
| DeepSeek MoE | 2 | 16.4B | 36.6 GB | RTX A6000 | $0.363/hr |
| DeepSeek VL | 3 | 2.0B to 7.3B | 4.4 GB | V100 | $0.088/hr |
| DeepSeek OCR | 2 | 3.3B to 3.4B | 7.5 GB | RTX 4070 Super | $0.121/hr |
Run and fine-tune DeepSeek
Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.
- DeepSeek V4: how to run, fine-tune
DeepSeek compared with other models
Side-by-side VRAM, context length and license for models of a similar size.
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen2.5-14B-Instruct
- DeepSeek-R1 vs GLM-5.2
- DeepSeek-R1 vs NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- DeepSeek-V2-Lite-Chat vs Qwen2.5-14B-Instruct
- DeepSeek-R1 vs Kimi-K2-Instruct
- DeepSeek-V2-Lite vs Qwen2.5-14B-Instruct
- DeepSeek-R1 vs Llama-3.1-405B
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen2.5-Coder-14B-Instruct
- DeepSeek-V3.2 vs GLM-5.2
- DeepSeek-V3.2 vs NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- DeepSeek-V2-Lite-Chat vs Qwen2.5-Coder-14B-Instruct
- DeepSeek-V3.2 vs Kimi-K2-Instruct
- DeepSeek-V2-Lite vs Qwen2.5-Coder-14B-Instruct
- DeepSeek-V3.2 vs Llama-3.1-405B
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen3-14B
- DeepSeek-V3-0324 vs GLM-5.2
- DeepSeek-V3 vs GLM-5.2
- DeepSeek-V2-Lite-Chat vs Qwen3-14B
- DeepSeek-R1-0528-Qwen3-8B vs Llama-3.1-8B-Instruct
- DeepSeek-V2-Lite vs Qwen3-14B
- DeepSeek-V3.1 vs GLM-5.2
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen3-14B-Base
- DeepSeek-V3-0324 vs NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- DeepSeek-Coder-V2-Lite-Instruct vs phi-4
- DeepSeek-V3-0324 vs Kimi-K2-Instruct
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen1.5-MoE-A2.7B
- DeepSeek-V3-0324 vs Llama-3.1-405B
- DeepSeek-R1-0528-Qwen3-8B vs Qwen2.5-Coder-7B-Instruct
- DeepSeek-Coder-V2-Lite-Instruct vs Qwen2.5-Coder-14B
- DeepSeek-V3-0324 vs GLM-5.3
- DeepSeek-V3 vs NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- DeepSeek-V2-Lite-Chat vs Qwen3-14B-Base
- DeepSeek-V3 vs Kimi-K2-Instruct
- DeepSeek-R1-0528-Qwen3-8B vs Meta-Llama-3-8B-Instruct
- DeepSeek-V2-Lite vs Qwen3-14B-Base
- DeepSeek-V2 vs GLM-4.5
- DeepSeek-V3 vs Llama-3.1-405B
Best models by task
DeepSeek models appear on these ranked task pages.
Sources
- https://huggingface.co/api/models?author=deepseek-ai&sort=createdAt&direction=-1&limit=50
- https://huggingface.co/deepseek-ai/DeepSeek-V2
- https://huggingface.co/api/models/deepseek-ai/DeepSeek-V2
- https://huggingface.co/deepseek-ai/DeepSeek-V3
- https://huggingface.co/deepseek-ai/DeepSeek-V3.1
- https://huggingface.co/deepseek-ai/DeepSeek-V3.2
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash
- https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- https://huggingface.co/deepseek-ai/DeepSeek-R1
- https://huggingface.co/deepseek-ai/DeepSeek-OCR
Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.