How to run Mistral Small 3.2 with vLLM

vLLM version, Mistral tokenizer flags, system prompt, temperature, image input and tool calling for Mistral Small 3.2.

This guide covers serving the Mistral Small family, based on the Mistral-Small-3.2-24B-Instruct-2506 card. Sizing and launch commands per model appear below.

Engine support

  • vLLM is the recommended engine. The card says to use vLLM 0.9.1 or newer, which should install mistral_common 1.6.2 or newer.
  • Transformers works with mistral-common 1.6.2 or newer, using MistralTokenizer and Mistral3ForConditionalGeneration.
  • The card does not mention llama.cpp, Ollama or SGLang.

Serving flags

The card's vLLM command uses the Mistral-native formats:

  • --tokenizer_mode mistral
  • --config_format mistral
  • --load_format mistral
  • --tool-call-parser mistral with --enable-auto-tool-choice for function calling
  • --limit-mm-per-prompt to allow images (the card sets 10 per prompt)

Prompting

  • Use a low temperature. The card recommends 0.15.
  • Load the system prompt from SYSTEM_PROMPT.txt in the model repository, applying the date formatting shown in the usage examples.
  • The model accepts image inputs. The card's examples use a maximum of 131072 tokens.

Pitfalls

  • Use all three mistral format flags together, as in the card's command.
  • Tool calling needs both the parser flag and auto tool choice.

Launch commands by Mistral Small size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

No Mistral Small model of a size we can compute is in the catalog yet. See the Mistral Small model list for what is published.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.