What GPU do I need to run MiniMaxAI/MiniMax-M2.5?
A 229B (MoE) language model for chat and instruction-following. 228.7B parameters, published in F8_E4M3. View on Hugging Face
MiniMax-M2.5 is published by MiniMaxAI on Hugging Face, with 487,430 downloads and 1,507 likes to date. It's a MiniMaxM2ForCausalLM model built for text-generation, published natively in F8_E4M3.
What MiniMax-M2.5 is
MiniMax-M2.5 is a 229B-parameter mixture-of-experts language model published by MiniMax on Hugging Face. It is released under Custom license.
License note: a lab-specific license (tagged "other" on Hugging Face); read the model card's own license section before commercial use. Facts in this section are sourced from MiniMax-M2.5's Hugging Face model card, not benchmarked by Aquanode.
What it's used for
- Chat assistants
- Instruction following
- Synthetic data generation
VRAM required & cheapest live GPU fit
Required VRAM = weight size at each precision, plus a fixed overhead for KV-cache, activations, and fragmentation. Full formula and assumptions: methodology.
| Precision | Weight size | Required VRAM | Cheapest live fit | GPUs needed | Est. $/hr (full fit) |
|---|---|---|---|---|---|
| FP8 (native) | 213.0 GB | 255.6 GB | RTX 4080 Super | 8 | $2.71/hr |
| INT4 (quantized) | 106.5 GB | 127.8 GB | RTX 5060 Ti | 8 | $0.880/hr |
A GPU is only matched to a row if its hardware supports that precision, and the primary recommendation is always a single-GPU fit when one exists.
INT4 requires a quantized checkpoint actually published for this model, check its Hugging Face page before relying on this row.
Cheapest way to run MiniMax-M2.5 at its published (F8_E4M3) precision: 8× RTX 4080 Super, at $0.338/hr per GPU ($2.71/hr total). Quantizing to FP8 or INT4 (rows above) can cost less, but requires a compatible quantized checkpoint to exist for this model.
MiniMax-M2.5: common questions
Can MiniMax-M2.5 run on a single GPU?
No. At FP8 (native) it needs 255.6 GB of VRAM, more than a single 32 GB desktop card holds. The cheapest capable card in the live feed is a 32.0 GB RTX 4080 Super, and it takes 8 of them.
Is MiniMax-M2.5 already quantized?
Yes. It is published in FP8, one byte per parameter, so the 255.6 GB figure is already a quantized footprint rather than a full-precision one. Only INT4 goes below it, at 127.8 GB. A GPU without FP8 tensor cores cannot run it as published, which is why cards here are matched on precision support and not on VRAM alone.
How many GPUs do I need to run MiniMax-M2.5?
8 at FP8 (native). It needs 255.6 GB of VRAM and the cheapest capable live offer is a 32.0 GB RTX 4080 Super, so 8 of them come to $2.71/hr in total.
Does quantizing MiniMax-M2.5 lower the GPU bill?
Yes. At FP8 (native) the cheapest live fit is 8 RTX 4080 Super cards at $2.71/hr. At INT4 (quantized) it drops to 8 RTX 5060 Ti cards at $0.880/hr, provided a quantized checkpoint exists for it.
How to run MiniMax-M2.5
Run MiniMax-M2.5 with vLLM
Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
vllm serve MiniMaxAI/MiniMax-M2.5 --tensor-parallel-size 8Run MiniMax-M2.5 with GGUF quantizations
Prebuilt GGUF weights published at lmstudio-community/MiniMax-M2.5-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio.
llama-server -hf lmstudio-community/MiniMax-M2.5-GGUFSource: https://huggingface.co/lmstudio-community/MiniMax-M2.5-GGUF
Deploy MiniMax-M2.5 on Aquanode
Aquanode has no one-click deploy template for MiniMax-M2.5; you install the inference engine yourself with the commands below. Aquanode sells GPU pods billed per second, not a hosted inference API.
- Launch a bare GPU pod sized to the requirement above (8× RTX 4080 Super or larger).
- Open a terminal on the pod, or save one of the commands above as a startup script so it runs automatically the first time the pod boots.
- Run the command and connect to the resulting endpoint.
Weight-to-VRAM math, the fit rules, and how live prices are normalized: full methodology.
More MiniMax M2 models
- MiniMax-M2.7 (228.7B, F8_E4M3)
- MiniMax-M2 (228.7B, F8_E4M3)
- MiniMax-M2.1 (228.7B, F8_E4M3)
Related reading: H100 pricing and specs, The best GPUs for AI, ranked, and Best GPU for LLM inference.