How to fine-tune MiniMax-M3 and its license limits

Fine-tuning MiniMax-M3: only a post-trained checkpoint is listed, thinking modes, trust-remote-code, and the attribution and notice terms.

Documentation for adapting the MiniMax-M3 family is thin. The card names unsloth as a lightweight fine-tuning option and gives no recipe. See fine-tuning and the training cost calculator for method and cost choices.

Checkpoint

The card lists a single post-trained model (about 428B parameters, about 23B active) and no base checkpoint. The Hub shows community fine-tunes and adapters, so LoRA on the released model is the path the ecosystem uses. The card gives no LoRA target modules.

Loading

trust_remote_code=True is required. Transformers loads it through AutoProcessor and AutoModelForMultimodalLM, so the model has image and video inputs. Decide whether your data is text only.

Chat format

The template ships with the model. Reasoning is controlled by thinking: enabled, adaptive or disabled. Match the mode in training data to the mode you will serve with. The card recommends temperature=1.0 and top_p=0.95 for inference.

License

The license is the custom minimax-community license. Per the LICENSE file:

  • Non-commercial use needs no attribution or authorization.
  • Commercial use includes deploying a model that has been fine-tuned, instruction-tuned or otherwise modified.
  • Commercial derivatives must prominently display "Built with MiniMax M3".
  • Under $20M yearly revenue, send a one-time notice to api@minimax.io. Above $20M, obtain prior written authorization from MiniMax.

The license also lists prohibited uses. Read the full text before releasing a derivative.

Memory to fine-tune MiniMax M3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

No MiniMax M3 model of a size we can compute is in the catalog yet. See the MiniMax M3 model list for what is published.

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.