MiniMax models, generation by generation

MiniMax's open-weight line from Text-01 to M2.x and M3: sizes, active parameters, context, licenses, and which generation to use today.

What the MiniMax line is

MiniMax publishes open weights under the MiniMaxAI organization on Hugging Face. The language models are all mixture-of-experts (MoE) designs, and each of the three generations makes a different bet. Text-01 bets on very long context through a hybrid of linear and softmax attention. The M2 series bets on small active size for coding agents. M3 returns to long context with sparse attention and adds native image and video input. The organization also publishes non-text systems, covered at the end.

Licensing varies by generation and none of the generations below has a plain Apache 2.0 or MIT model license, so check the license file for your use case.

Generations in order

MiniMax-Text-01 (the Hugging Face repository was created June 2025). Text-01 has 456 billion total parameters, of which 45.9 billion are active per token. It has 80 layers, and its card describes a hybrid attention layout in which one softmax attention layer follows every seven lightning-attention layers, with 32 experts. Lightning attention is a linear-attention variant, which is what lets the context grow. The card says training context was extended to 1 million tokens and that the model can handle up to 4 million tokens at inference. The model license is a custom MiniMax model agreement, with the code under MIT.

MiniMax-M2 (October 2025, repository date). M2 is a much smaller MoE: 230 billion total parameters and 10 billion active. The card positions it for coding and agentic tool use, and argues that a 10 billion active size keeps the plan, act and verify loop responsive and cheaper to batch. Note that total size still determines how much memory you need, not the active count. The card tags the license as modified MIT.

The same hub holds the point releases, all in the same size class (their Hugging Face metadata lists about 229 billion parameters, essentially unchanged from M2):

  • M2.1 (December 2025).
  • M2.5 (February 2026). The card describes extensive reinforcement-learning training in real-world environments and says it was released in two versions, M2.5 and M2.5-Lightning, that are identical in capability and differ in speed. Modified MIT.
  • M2.7 (April 2026). The card calls it the first model deeply participating in its own evolution, and highlights building agent harnesses and productivity tasks. Its license is a separate other license, so read the file rather than assuming M2's terms carry over.

Benchmark figures on these cards are vendor-reported.

MiniMax-M3 (June 2026, repository date). M3 is a native multimodal model with a 1 million token context window, about 428 billion total and 23 billion active parameters. Its card describes training on text, image and video from the first step, and a new attention operator, MiniMax Sparse Attention (MSA), which MiniMax says gives 9x faster prefill and 15x faster decode than M2 at 1 million tokens context (vendor-reported). It supports three reasoning modes (enabled, adaptive, disabled) and lists SGLang, vLLM, Transformers, KTransformers and unsloth as serving options. The license is minimax-community; read it before commercial use. An MXFP8 variant is also on the hub.

Other MiniMax releases

MiniMax-H3, a general-purpose omni-modal generative system that understands text, image, video and audio and generates video with stereo audio up to 2K, and MiniMax-Music3 are published by the same organization (July and August 2026 repositories). They are not language models in this line, and H3 has its own community license agreement.

Which generation to use today

  • A coding or agent model with the smallest active size: M2.5 or M2.7, at roughly 10 billion active parameters (M2's card figure; the point releases list the same tensor total). M2.5 is modified MIT; M2.7 has its own license.
  • Images or video as input, or contexts approaching 1 million tokens: M3. It is the only MiniMax language model here with multimodal input and the 1 million token window stated on its card.
  • Studying or reproducing the linear-attention hybrid: Text-01. It is the oldest generation and its card lists text input only.
  • Lowest license friction: none of the three is Apache 2.0 or MIT. M2 and M2.5 are modified MIT, the closest to it.

All of these can run on Aquanode GPUs.

MiniMax generations

Every MiniMax generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
MiniMax-Text1456.1B1019 GBNo live fit–
MiniMax M24228.7B to 228.7B256 GBRTX 4080 Super × 8$2.71/hr
MiniMax M343.1B to 440.3B955 GBNo live fit–

Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.

Other MiniMax lines

Side lines from MiniMax: specialised or companion models outside the main MiniMax generations.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
MiniMax H3333.1B to 35.0B74.0 GBA100$1.21/hr
MiniMax Music312.4B10.9 GBV100$0.088/hr

Run and fine-tune MiniMax

Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.

MiniMax compared with other models

Side-by-side VRAM, context length and license for models of a similar size.

Best models by task

MiniMax models appear on these ranked task pages.

Sources

Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.