How to fine-tune Mistral Small (LoRA, QLoRA)

Start from the Mistral Small base or instruct checkpoint, with Apache 2.0 licensing, Tekken tokenizer and vision-tower notes.

This guide covers adapting the Mistral Small family. See /fine-tuning and the training cost calculator for planning.

Which checkpoint

  • Base: Mistral-Small-3.1-24B-Base-2503 is a pretrained model that is "not ready to work as an instruction model out-of-the-box" per its card. Choose it when you want to teach your own instruction style.
  • Instruct: the 3.2 Instruct 2506 model is built on that base. Start here for LoRA or QLoRA on a specific task.

License

Both cards list the Apache 2.0 license. The base card describes it as allowing use and modification for commercial and non-commercial purposes.

Tokenizer and format

  • The model uses the Tekken tokenizer with a 131k vocabulary.
  • The Instruct card relies on mistral-common 1.6.2 or newer for tokenization, so build training examples with the same tokenizer rather than a generic template.
  • The Instruct card recommends a system prompt loaded from SYSTEM_PROMPT.txt. Decide whether your training data keeps or replaces it.

Multimodal

The base model includes vision understanding, and the Instruct model takes image inputs.

Not documented

The cards give no recommended LoRA target modules or training recipe. The base card says Transformers weights were only "vibe-checked" and recommends vLLM for reliable behavior, so validate your adapter in the engine you serve with.

Memory to fine-tune Mistral Small, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

No Mistral Small model of a size we can compute is in the catalog yet. See the Mistral Small model list for what is published.

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.