Mistral models, generation by generation
Mistral AI's open-weight line from Mistral 7B to Mistral Small 4 and Large 3: sizes, dense vs MoE, context length, licenses, and which to run today.
What the Mistral line is
Mistral AI publishes open-weight language models under the mistralai organization on Hugging Face. The line started with a small dense 7B model and has since grown into mixture-of-experts (MoE) models, multimodal models, and a speech model. Most of the open releases use the Apache 2.0 license, which allows commercial use without a revenue cap. The exceptions are called out below.
Generations in order
Mistral 7B (2023). The first release, Mistral-7B-v0.1, is a 7 billion parameter dense transformer. Its model card lists grouped-query attention, sliding-window attention and a byte-fallback BPE tokenizer, and the license is Apache 2.0. The card states that it outperforms Llama 2 13B on the benchmarks Mistral tested. Many community fine-tunes (Zephyr, OpenChat and others) start from this base, so its hub is large.
Mixtral (December 2023). Mixtral-8x7B is a sparse mixture-of-experts model with 47 billion total parameters, so not all of them are active for each token. It keeps the Apache 2.0 license. Mistral states that it outperforms Llama 2 70B on most of the benchmarks it tested. The memory you need follows the 47 billion total parameters, not the smaller active count, which is the usual surprise with MoE models.
Mistral Small. Mistral-Small-24B-Instruct-2501 (January 2025, "Mistral Small 3") is a 24 billion parameter dense model with a 32,000 token context window, native function calling and JSON output, and Apache 2.0 licensing. It uses the Tekken tokenizer with a 131,000 entry vocabulary.
Mistral Small 4 (model id 2603) changes the design. It is a MoE model with 119 billion total parameters and 6.5 billion active per token, drawn from 128 experts with 4 active. It accepts text and images, has a 256,000 token context window, supports a per-request reasoning setting, and is Apache 2.0. Its card says it unifies three previous model families into one set of weights.
Ministral. Mistral AI's own edge line is Ministral 3, released December 2, 2025, in 3B, 8B and 14B sizes, each as Base, Instruct and Reasoning variants. The card lists vision input, a 256,000 token context window and Apache 2.0. A caution about the hub: the model it currently holds is Ministral-3b-instruct, a 3B English chat model published by an organization named ministral, not by Mistral AI. Check the publisher before assuming it is the Mistral AI edge line.
Large and Medium, without a hub. Mistral-Large-3-675B-Instruct-2512 (December 2025) is a multimodal MoE with 675 billion total and 41 billion active parameters, a 256,000 token context window, and Apache 2.0. Mistral-Medium-3.5-128B is a 128 billion parameter dense model with text and image input, a 256,000 token context window, and a modified MIT license that carries restrictions for large-revenue companies. Read the license text before using it commercially.
Side lines
- Voxtral is the speech line. Voxtral Mini 4B Realtime 2602 is a streaming speech-to-text model of about 4 billion parameters (roughly 3.4B language model plus a 970M audio encoder), covering 13 languages under Apache 2.0, with configurable transcription delay from 80 ms to 2.4 seconds.
- Mistral experimental holds models published under the
mistral-experimentalorganization, such as Pixtral 12B, a vision-language model (about 13 billion parameters in BF16, Apache 2.0).
Which generation to use today
- A fully permissive license, mid-size, text only: Mistral Small 3 (24B dense). Dense weights are simple to serve, and 32,000 tokens is enough for most chat and tool-calling work.
- Long context, images and a reasoning toggle in one model: Mistral Small 4. It trades a larger memory footprint (119 billion parameters must be resident) for 256,000 tokens of context and only 6.5 billion active parameters per token.
- Small footprint or edge use: Ministral 3 in the 3B or 8B size, which the card says can run on as little as 12 GB of RAM when quantized.
- A frontier-size open model with Apache 2.0: Mistral Large 3, if you have multi-GPU capacity for 675 billion parameters.
- Fine-tuning a small base: Mistral 7B has the largest set of community derivatives on its hub, but it is the oldest line here and its card lists no vision input or long-context setting.
Any of these can run on Aquanode GPUs.
Mistral generations
Every Mistral generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| Mistral 7B | 7 | 7.2B to 7.2B | 16.2 GB | RTX A5000 | $0.176/hr |
| Mixtral | 1 | 46.7B | 104 GB | RTX A5000 × 5 | $0.880/hr |
| Mistral Small | 3 | 23.6B to 24.0B | 52.7 GB | A100 | $1.21/hr |
| ministral | 1 | 3.3B | 7.4 GB | RTX 4070 Super | $0.121/hr |
Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.
Other Mistral lines
Side lines from Mistral AI: specialised or companion models outside the main Mistral generations.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| Voxtral | 1 | 4.4B | 9.9 GB | RTX 4070 Super | $0.121/hr |
| mistral-experimental | 1 | 12.7B | 28.3 GB | RTX A6000 | $0.363/hr |
Run and fine-tune Mistral
Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.
- Mistral Small: how to run, fine-tune
Mistral compared with other models
Side-by-side VRAM, context length and license for models of a similar size.
Best models by task
Mistral models appear on these ranked task pages.
Sources
- https://huggingface.co/mistralai/Mistral-7B-v0.1
- https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1
- https://huggingface.co/mistralai/Mistral-Small-24B-Instruct-2501
- https://huggingface.co/mistralai/Mistral-Small-4-119B-2603
- https://huggingface.co/mistralai/Mistral-Medium-3.5-128B
- https://huggingface.co/mistralai/Ministral-3-8B-Reasoning-2512
- https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512
- https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602
- https://huggingface.co/ministral/Ministral-3b-instruct
- https://huggingface.co/mistral-experimental/pixtral-12b
Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.