DeepSeek-Coder-V2-Lite-Instruct vs Qwen1.5-MoE-A2.7B
DeepSeek-Coder-V2-Lite-Instruct (15.7B parameters) and Qwen1.5-MoE-A2.7B (14.3B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.
Side by side
| Fact | DeepSeek-Coder-V2-Lite-Instruct | Qwen1.5-MoE-A2.7B |
|---|---|---|
| Parameters | 15.7B | 14.3B |
| Architecture | Multi-head latent attention; mixture of 64 experts, 6 active per token | Multi-head attention; mixture of 60 experts, 4 active per token |
| Context length | 163,840 tokens | 8,192 tokens |
| License | – | – |
| Published precision | BF16 | BF16 |
| VRAM needed, As published | 35.1 GB | 32.0 GB |
| VRAM needed, FP8 | 17.6 GB | 16.0 GB |
| VRAM needed, INT4 | 8.8 GB | 8.0 GB |
| Cheapest live fit, As published | RTX A6000 · $0.363/hr | RTX 4080 Super · $0.338/hr |
| Cheapest live fit, FP8 | RTX 4000 SFF Ada · $0.198/hr | RTX 4000 SFF Ada · $0.198/hr |
| Cheapest live fit, INT4 | RTX 4070 Super · $0.121/hr | RTX 4070 Super · $0.121/hr |
| KV cache per token (16-bit) | 30 KB | 192 KB |
| KV cache at 32k tokens | 0.95 GB | 6.00 GB |
| KV cache at 128k tokens | 3.80 GB | 24.0 GB |
VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.
Which to pick
- Qwen1.5-MoE-A2.7B needs less VRAM at its published precision (32.0 GB against 35.1 GB), so it fits on a smaller GPU.
- DeepSeek-Coder-V2-Lite-Instruct lists the longer context window (163,840 tokens against 8,192).
- DeepSeek-Coder-V2-Lite-Instruct caches less per sequence at 32k tokens (0.9 GB against 6.0 GB), leaving more memory for batching.
- Qwen1.5-MoE-A2.7B has the cheaper live GPU fit at its published precision ($0.338/hr against $0.363/hr).
These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.
Keep reading
- DeepSeek-Coder-V2-Lite-Instruct: full VRAM table and live GPU fit
- Qwen1.5-MoE-A2.7B: full VRAM table and live GPU fit
- The DeepSeek model series
- The Qwen model series
- All models that fit in 48 GB
- All models that fit in 32 GB
Other comparisons