DeepSeek-V3 vs Kimi-K2-Instruct
DeepSeek-V3 (684.5B parameters) and Kimi-K2-Instruct (1026.4B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.
Side by side
| Fact | DeepSeek-V3 | Kimi-K2-Instruct |
|---|---|---|
| Parameters | 684.5B (Mixture-of-experts: 8 of 256 experts active per token (exact active-parameter count not stated on the model card)) | 1026.4B (Mixture-of-experts: 8 of 384 experts active per token (exact active-parameter count not stated on the model card)) |
| Architecture | Multi-head latent attention; mixture of 256 experts, 8 active per token | Multi-head latent attention; mixture of 384 experts, 8 active per token |
| Context length | 160K tokens (163,840) | 128K tokens (131,072) |
| License | Not stated | Custom license |
| Published precision | F8_E4M3 | F8_E4M3 |
| VRAM needed, As published | 765 GB | 1147 GB |
| VRAM needed, FP8 | Not a smaller option | Not a smaller option |
| VRAM needed, INT4 | 383 GB | 574 GB |
| Cheapest live fit, As published | RTX PRO 6000 × 8 · $11.75/hr | No live fit |
| Cheapest live fit, FP8 | – | – |
| Cheapest live fit, INT4 | RTX A6000 × 8 · $2.90/hr | RTX PRO 6000 × 6 · $8.82/hr |
| KV cache per token (16-bit) | 69 KB | 69 KB |
| KV cache at 32k tokens | 2.14 GB | 2.14 GB |
| KV cache at 128k tokens | 8.58 GB | 8.58 GB |
VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.
Which to pick
- DeepSeek-V3 needs less VRAM at its published precision (765 GB against 1147 GB), so it fits on a smaller GPU.
- DeepSeek-V3 lists the longer context window (163,840 tokens against 131,072).
- Licenses differ: DeepSeek-V3 is under Not stated and Kimi-K2-Instruct under Custom license. Read both before commercial use.
These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.
Keep reading
- DeepSeek-V3: full VRAM table and live GPU fit
- Kimi-K2-Instruct: full VRAM table and live GPU fit
- The DeepSeek model series
- The Kimi model series
Other comparisons