How to run MiniMax-M3 with vLLM and SGLang

Serving MiniMax-M3: vLLM version, parser flags, block size, the three thinking modes, trust-remote-code and image and video inputs.

This guide covers serving the MiniMax-M3 family. Memory sizing and launch commands appear below.

Engine support

The model card points to SGLang, vLLM, Transformers, KTransformers, unsloth and ATOM (AMD ROCm), and lists quantized builds for llama.cpp, Ollama and LM Studio. It states no minimum version. The vLLM recipe says vLLM 0.24.0 or newer is required and notes that support had not yet shipped in a stable release when it was written, so check which build you install.

vLLM flags

  • --tool-call-parser minimax_m3 and --reasoning-parser minimax_m3 (earlier MiniMax releases used minimax_m2, so do not reuse old commands)
  • --block-size 128, called mandatory on all platforms for the sparse attention indexing
  • --trust-remote-code, listed as required for the NVFP4 and MXFP4 variants; the card also requires trust_remote_code=True for Transformers
  • --mm-encoder-tp-mode data for data-parallel vision encoding at high tensor parallelism
  • --language-model-only for text-only workloads

Thinking modes

Set thinking_mode through chat_template_kwargs:

  • enabled: always reasons
  • adaptive: the model decides (the recipe's default)
  • disabled: no reasoning, lowest latency

Context and inputs

Context is 1M tokens. The card documents no rope or YaRN override. The model accepts text, images and video natively.

Sampling

The card recommends temperature=1.0 and top_p=0.95.

Pitfalls

  • Missing --block-size 128.
  • Using the old minimax_m2 parsers.
  • Loading the vision stack when you only need text.

Launch commands by MiniMax M3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

No MiniMax M3 model of a size we can compute is in the catalog yet. See the MiniMax M3 model list for what is published.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.