MiniMax models, generation by generation
MiniMax's open-weight line from Text-01 to M2.x and M3: sizes, active parameters, context, licenses, and which generation to use today.
What the MiniMax line is
MiniMax publishes open weights under the MiniMaxAI organization on Hugging Face. The language models are all mixture-of-experts (MoE) designs, and each of the three generations makes a different bet. Text-01 bets on very long context through a hybrid of linear and softmax attention. The M2 series bets on small active size for coding agents. M3 returns to long context with sparse attention and adds native image and video input. The organization also publishes non-text systems, covered at the end.
Licensing varies by generation and none of the generations below has a plain Apache 2.0 or MIT model license, so check the license file for your use case.
Generations in order
MiniMax-Text-01 (the Hugging Face repository was created June 2025). Text-01 has 456 billion total parameters, of which 45.9 billion are active per token. It has 80 layers, and its card describes a hybrid attention layout in which one softmax attention layer follows every seven lightning-attention layers, with 32 experts. Lightning attention is a linear-attention variant, which is what lets the context grow. The card says training context was extended to 1 million tokens and that the model can handle up to 4 million tokens at inference. The model license is a custom MiniMax model agreement, with the code under MIT.
MiniMax-M2 (October 2025, repository date). M2 is a much smaller MoE: 230 billion total parameters and 10 billion active. The card positions it for coding and agentic tool use, and argues that a 10 billion active size keeps the plan, act and verify loop responsive and cheaper to batch. Note that total size still determines how much memory you need, not the active count. The card tags the license as modified MIT.
The same hub holds the point releases, all in the same size class (their Hugging Face metadata lists about 229 billion parameters, essentially unchanged from M2):
- M2.1 (December 2025).
- M2.5 (February 2026). The card describes extensive reinforcement-learning training in real-world environments and says it was released in two versions, M2.5 and M2.5-Lightning, that are identical in capability and differ in speed. Modified MIT.
- M2.7 (April 2026). The card calls it the first model deeply participating in its own evolution, and highlights building agent harnesses and productivity tasks. Its license is a separate
otherlicense, so read the file rather than assuming M2's terms carry over.
Benchmark figures on these cards are vendor-reported.
MiniMax-M3 (June 2026, repository date). M3 is a native multimodal model with a 1 million token context window, about 428 billion total and 23 billion active parameters. Its card describes training on text, image and video from the first step, and a new attention operator, MiniMax Sparse Attention (MSA), which MiniMax says gives 9x faster prefill and 15x faster decode than M2 at 1 million tokens context (vendor-reported). It supports three reasoning modes (enabled, adaptive, disabled) and lists SGLang, vLLM, Transformers, KTransformers and unsloth as serving options. The license is minimax-community; read it before commercial use. An MXFP8 variant is also on the hub.
Other MiniMax releases
MiniMax-H3, a general-purpose omni-modal generative system that understands text, image, video and audio and generates video with stereo audio up to 2K, and MiniMax-Music3 are published by the same organization (July and August 2026 repositories). They are not language models in this line, and H3 has its own community license agreement.
Which generation to use today
- A coding or agent model with the smallest active size: M2.5 or M2.7, at roughly 10 billion active parameters (M2's card figure; the point releases list the same tensor total). M2.5 is modified MIT; M2.7 has its own license.
- Images or video as input, or contexts approaching 1 million tokens: M3. It is the only MiniMax language model here with multimodal input and the 1 million token window stated on its card.
- Studying or reproducing the linear-attention hybrid: Text-01. It is the oldest generation and its card lists text input only.
- Lowest license friction: none of the three is Apache 2.0 or MIT. M2 and M2.5 are modified MIT, the closest to it.
All of these can run on Aquanode GPUs.
MiniMax generations
Every MiniMax generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| MiniMax-Text | 1 | 456.1B | 1019 GB | No live fit | – |
| MiniMax M2 | 4 | 228.7B to 228.7B | 256 GB | RTX 4080 Super × 8 | $2.71/hr |
| MiniMax M3 | 4 | 3.1B to 440.3B | 955 GB | No live fit | – |
Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.
Other MiniMax lines
Side lines from MiniMax: specialised or companion models outside the main MiniMax generations.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| MiniMax H3 | 3 | 33.1B to 35.0B | 74.0 GB | A100 | $1.21/hr |
| MiniMax Music3 | 1 | 2.4B | 10.9 GB | V100 | $0.088/hr |
Run and fine-tune MiniMax
Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.
- MiniMax M3: how to run, fine-tune
MiniMax compared with other models
Side-by-side VRAM, context length and license for models of a similar size.
- MiniMax-M2.7 vs Qwen3-235B-A22B
- GLM-4.5 vs MiniMax-M2.7
- MiniMax-M2.7 vs Qwen3-235B-A22B-Instruct-2507
- GLM-4.7 vs MiniMax-M2.7
- MiniMax-M2.5 vs Qwen3-235B-A22B
- GLM-4.5 vs MiniMax-M2.5
- MiniMax-M2.5 vs Qwen3-235B-A22B-Instruct-2507
- GLM-4.7 vs MiniMax-M2.5
- MiniMax-M2 vs Qwen3-235B-A22B
- GLM-4.5 vs MiniMax-M2
- MiniMax-M2 vs Qwen3-235B-A22B-Instruct-2507
- GLM-4.7 vs MiniMax-M2
- Hy3-preview vs MiniMax-M2
Best models by task
MiniMax models appear on these ranked task pages.
Sources
- https://huggingface.co/MiniMaxAI/MiniMax-Text-01-hf
- https://huggingface.co/MiniMaxAI/MiniMax-M2
- https://huggingface.co/MiniMaxAI/MiniMax-M2.5
- https://huggingface.co/MiniMaxAI/MiniMax-M2.7
- https://huggingface.co/MiniMaxAI/MiniMax-M3
- https://huggingface.co/MiniMaxAI/MiniMax-H3
- https://huggingface.co/api/models/MiniMaxAI/MiniMax-Text-01-hf
- https://huggingface.co/api/models/MiniMaxAI/MiniMax-M2.1
- https://huggingface.co/api/models/MiniMaxAI/MiniMax-M3
Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.