MiniMax-M2 vs Qwen3-235B-A22B-Instruct-2507
MiniMax-M2 (228.7B parameters) and Qwen3-235B-A22B-Instruct-2507 (235.1B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.
Side by side
| Fact | MiniMax-M2 | Qwen3-235B-A22B-Instruct-2507 |
|---|---|---|
| Parameters | 228.7B (Mixture-of-experts: 8 of 256 experts active per token (10B active parameters, stated on the model card)) | 235.1B |
| Architecture | Grouped-query attention; mixture of 256 experts, 8 active per token | Grouped-query attention; mixture of 128 experts, 8 active per token |
| Context length | 192K tokens (196,608) | 262,144 tokens |
| License | Modified MIT | – |
| Published precision | F8_E4M3 | BF16 |
| VRAM needed, As published | 256 GB | 525 GB |
| VRAM needed, FP8 | Not a smaller option | 263 GB |
| VRAM needed, INT4 | 128 GB | 131 GB |
| Cheapest live fit, As published | RTX 4080 Super × 8 · $2.71/hr | RTX PRO 6000 × 6 · $8.25/hr |
| Cheapest live fit, FP8 | – | RTX PRO 6000 × 3 · $4.13/hr |
| Cheapest live fit, INT4 | RTX A5000 × 6 · $1.06/hr | RTX A5000 × 6 · $1.06/hr |
| KV cache per token (16-bit) | 248 KB | 188 KB |
| KV cache at 32k tokens | 7.75 GB | 5.88 GB |
| KV cache at 128k tokens | 31.0 GB | 23.5 GB |
VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.
Which to pick
- MiniMax-M2 needs less VRAM at its published precision (256 GB against 525 GB), so it fits on a smaller GPU.
- Qwen3-235B-A22B-Instruct-2507 lists the longer context window (262,144 tokens against 196,608).
- Qwen3-235B-A22B-Instruct-2507 caches less per sequence at 32k tokens (5.9 GB against 7.8 GB), leaving more memory for batching.
- MiniMax-M2 has the cheaper live GPU fit at its published precision ($2.71/hr against $8.25/hr).
These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.
Keep reading
- MiniMax-M2: full VRAM table and live GPU fit
- Qwen3-235B-A22B-Instruct-2507: full VRAM table and live GPU fit
- The MiniMax model series
- The Qwen model series
- All models that fit in 288 GB
Other comparisons