What GPU do I need to run MiniMaxAI/MiniMax-M2?
228.7B parameters, published in F8_E4M3. View on Hugging Face Full specs & deploy guide
MiniMax-M2 is published by MiniMaxAI on Hugging Face, with 258,299 downloads and 1,503 likes to date. It's a MiniMaxM2ForCausalLM model built for text-generation, published natively in F8_E4M3.
VRAM required & cheapest live GPU fit
Required VRAM = weight size at each precision, plus a fixed overhead for KV-cache, activations, and fragmentation. Full formula and assumptions: methodology.
| Precision | Weight size | Required VRAM | Cheapest live fit | GPUs needed | Est. $/hr (full fit) |
|---|---|---|---|---|---|
| FP8 (native) | 213.0 GB | 255.6 GB | RTX 4080 Super | 8 | $2.71/hr |
| INT4 (quantized) | 106.5 GB | 127.8 GB | RTX A5000 | 6 | $1.06/hr |
A GPU is only matched to a row if its hardware supports that precision, and the primary recommendation is always a single-GPU fit when one exists.
INT4 requires a quantized checkpoint actually published for this model, check its Hugging Face page before relying on this row.
Cheapest way to run MiniMax-M2 at its published (F8_E4M3) precision: 8× RTX 4080 Super, at $0.338/hr per GPU ($2.71/hr total). Quantizing to FP8 or INT4 (rows above) can cost less, but requires a compatible quantized checkpoint to exist for this model.
MiniMax-M2: common questions
Can MiniMax-M2 run on a single GPU?
No. At FP8 (native) it needs 255.6 GB of VRAM, more than a single 32 GB desktop card holds. The cheapest capable card in the live feed is a 32.0 GB RTX 4080 Super, and it takes 8 of them.
Is MiniMax-M2 already quantized?
Yes. It is published in FP8, one byte per parameter, so the 255.6 GB figure is already a quantized footprint rather than a full-precision one. Only INT4 goes below it, at 127.8 GB. A GPU without FP8 tensor cores cannot run it as published, which is why cards here are matched on precision support and not on VRAM alone.
How many GPUs do I need to run MiniMax-M2?
8 at FP8 (native). It needs 255.6 GB of VRAM and the cheapest capable live offer is a 32.0 GB RTX 4080 Super, so 8 of them come to $2.71/hr in total.
Does quantizing MiniMax-M2 lower the GPU bill?
Yes. At FP8 (native) the cheapest live fit is 8 RTX 4080 Super cards at $2.71/hr. At INT4 (quantized) it drops to 6 RTX A5000 cards at $1.06/hr, provided a quantized checkpoint exists for it.
Weight-to-VRAM math, the fit rules, and how live prices are normalized: full methodology.
More MiniMaxAI models
- MiniMax-H3 (33.1B, BF16)
- MiniMax-M2.7 (228.7B, F8_E4M3)
- MiniMax-M2.5 (228.7B, F8_E4M3)
- MiniMax-M3-MXFP8 (440.3B, F8_E4M3)
- MiniMax-M3 (427.0B, BF16)
- MiniMax-M2.1 (228.7B, F8_E4M3)
Related reading: H100 pricing and specs, The best GPUs for AI, ranked, and Best GPU for LLM inference.