How to run MiniMax-M3 with vLLM and SGLang
Serving MiniMax-M3: vLLM version, parser flags, block size, the three thinking modes, trust-remote-code and image and video inputs.
This guide covers serving the MiniMax-M3 family. Memory sizing and launch commands appear below.
Engine support
The model card points to SGLang, vLLM, Transformers, KTransformers, unsloth and ATOM (AMD ROCm), and lists quantized builds for llama.cpp, Ollama and LM Studio. It states no minimum version. The vLLM recipe says vLLM 0.24.0 or newer is required and notes that support had not yet shipped in a stable release when it was written, so check which build you install.
vLLM flags
--tool-call-parser minimax_m3and--reasoning-parser minimax_m3(earlier MiniMax releases usedminimax_m2, so do not reuse old commands)--block-size 128, called mandatory on all platforms for the sparse attention indexing--trust-remote-code, listed as required for the NVFP4 and MXFP4 variants; the card also requirestrust_remote_code=Truefor Transformers--mm-encoder-tp-mode datafor data-parallel vision encoding at high tensor parallelism--language-model-onlyfor text-only workloads
Thinking modes
Set thinking_mode through chat_template_kwargs:
enabled: always reasonsadaptive: the model decides (the recipe's default)disabled: no reasoning, lowest latency
Context and inputs
Context is 1M tokens. The card documents no rope or YaRN override. The model accepts text, images and video natively.
Sampling
The card recommends temperature=1.0 and top_p=0.95.
Pitfalls
- Missing
--block-size 128. - Using the old
minimax_m2parsers. - Loading the vision stack when you only need text.
Launch commands by MiniMax M3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
No MiniMax M3 model of a size we can compute is in the catalog yet. See the MiniMax M3 model list for what is published.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/MiniMaxAI/MiniMax-M3
- https://huggingface.co/MiniMaxAI/MiniMax-M3/raw/main/README.md
- https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3
Updated 2026-10-07.