Best open chat models
How to pick an open chat model: license, context, size and latency. Picks by GPU budget with facts taken from the publishers' model cards.
For a chat assistant you want an instruct-tuned model that follows instructions, holds a conversation, and answers fast. Pick on license, context length, and size, then test it on your own prompts. This page covers how to choose; the catalog below lists the instruct models.
What matters when picking a chat model
Instruct, not base. Base checkpoints continue text and do not reliably follow instructions. Chat apps want the instruct (or "it", "chat") variant from the same publisher.
Latency and concurrency. Chat is interactive, so time to first token and decode speed shape the experience. Smaller models, and mixture-of-experts models that run only some weights per token, respond faster. If many users share one GPU, the KV cache for each conversation also eats memory.
Context length. Long chats and document questions need room. Longer context costs memory, so match it to what your product really sends.
Tool calling and structured output. If the assistant calls functions or returns JSON, check that the card advertises native support.
License. Apache 2.0 and MIT cards carry no extra use restrictions of their own. Some chat models use custom community licenses with conditions, so read the card.
Precision. BF16 is the reference; a quantized build shrinks the footprint. Use the VRAM calculator to size it.
Picks by situation
Smallest and cheapest to serve: Phi-4-mini-instruct. Microsoft's card lists a 128K context and an MIT license, and the catalog shows about 3.8B parameters. It is a good fit for high-volume, low-latency assistants where a small model is enough.
Smaller GPUs (a 24 GB card): Mistral-Small-24B-Instruct-2501. The card lists 24B parameters, a 32k context window, and Apache 2.0, and highlights native function calling and JSON output. A 24B model in BF16 is a stretch for 24 GB, so expect to run a quantized build there or step up a card size.
One 80 GB GPU: gpt-oss-120b. OpenAI's card lists 117B parameters with 5.1B active and Apache 2.0, and says it fits a single 80GB GPU. Only a fraction of the weights run per token, which helps latency. It also has configurable reasoning effort, so you can keep chat replies quick or let it think longer on hard questions.
Lower latency from the same family: gpt-oss-20b. 21B parameters with 3.6B active, Apache 2.0, aimed at lower-latency and local use per the card.
Longest context: Phi-4-mini-instruct at 128K is the longest among these picks. For multimodal chat with even longer windows, see the vision-language page.
Most permissive license: Mistral Small and the gpt-oss models are Apache 2.0, and Phi-4-mini-instruct is MIT, all per their cards.
Serving notes
Run the model behind an inference server that batches requests, since chat traffic is bursty. Keep the system prompt short, because it is re-read on every turn. For a first deployment, pick the smallest model that passes your own evaluation set, then scale up only if it fails.
Aquanode rents GPUs by the hour, so you can try a 4B model and a 100B-class model on the right card sizes and keep the one that works. See pricing. Model-card claims are vendor-reported, so measure on your own traffic.
Open models for chat and assistants
The 60 most downloaded of 473 models in the catalog, grouped by size. The rest are listed on the models directory.
Under 3B parameters
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen3-8B | 8.2B | 18.3 GB | Apache 2.0 | RTX A5000 | $0.176/hr |
| Qwen2.5-7B-Instruct | 7.6B | 17.0 GB | Apache 2.0 | RTX A5000 | $0.176/hr |
| Qwen2.5-3B-Instruct | 3.1B | 6.9 GB | Custom license | RTX 4070 Super | $0.121/hr |
| Qwen3-4B | 4.0B | 9.0 GB | Apache 2.0 | RTX 4070 Super | $0.121/hr |
| Llama-3.1-8B-Instruct | 8.0B | 17.9 GB | Llama 3.1 Community License Agreement | RTX A5000 | $0.176/hr |
| Qwen3-4B-Instruct-2507 | 4.0B | 9.0 GB | – | RTX 4070 Super | $0.121/hr |
| Meta-Llama-3-8B-Instruct | 8.0B | 17.9 GB | – | RTX A5000 | $0.176/hr |
| Llama-3.2-3B-Instruct | 3.2B | 7.2 GB | Llama 3.2 Community License Agreement | RTX 4070 Super | $0.121/hr |
| Mistral-7B-Instruct-v0.2 | 7.2B | 16.2 GB | – | RTX A5000 | $0.176/hr |
| PowerMoE-3b | 3.4B | 15.1 GB | – | V100 | $0.088/hr |
| granite-4.1-8b | 8.8B | 19.7 GB | – | RTX A5000 | $0.176/hr |
| DeepSeek-R1-0528-Qwen3-8B | 8.2B | 18.3 GB | – | RTX A5000 | $0.176/hr |
| Llama-2-7b-hf | 6.7B | 15.1 GB | – | V100 | $0.088/hr |
| gemma-2-9b-it | 9.2B | 20.7 GB | – | RTX A5000 | $0.176/hr |
| Meta-Llama-3-8B | 8.0B | 17.9 GB | – | RTX A5000 | $0.176/hr |
| Apertus-8B-Instruct-2509 | 8.1B | 18.0 GB | – | RTX A5000 | $0.176/hr |
| Phi-3-mini-4k-instruct | 3.8B | 8.5 GB | MIT | RTX 4070 Super | $0.121/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen3-32B | 32.8B | 73.2 GB | Apache 2.0 | A100 | $1.21/hr |
| Qwen2.5-14B-Instruct | 14.8B | 33.0 GB | Apache 2.0 | RTX A6000 | $0.363/hr |
| Qwen3-30B-A3B | 30.5B | 68.2 GB | Apache 2.0 | A100 | $1.21/hr |
| Qwen2.5-32B-Instruct | 32.8B | 73.2 GB | – | A100 | $1.21/hr |
| GLM-4.7-Flash | 31.2B | 69.8 GB | MIT | A100 | $1.21/hr |
| Qwen3-14B | 14.8B | 33.0 GB | – | RTX A6000 | $0.363/hr |
| Qwen3-30B-A3B-Instruct-2507 | 30.5B | 68.2 GB | – | A100 | $1.21/hr |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | 31.6B | 70.6 GB | – | A100 | $1.21/hr |
| gpt-neox-20b | 20.7B | 46.4 GB | – | RTX A6000 | $0.363/hr |
| phi-4 | 14.7B | 32.8 GB | MIT | RTX A6000 | $0.363/hr |
40B to 150B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen-72B | 72.3B | 162 GB | – | RTX A5000 × 7 | $1.23/hr |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 123.6B | 276 GB | – | RTX A6000 × 6 | $2.18/hr |
150B parameters and up
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| DeepSeek-R1 | 684.5B | 765 GB | – | RTX PRO 6000 WS × 8 | $11.75/hr |
| DeepSeek-V3.2 | 685.4B | 766 GB | – | RTX PRO 6000 WS × 8 | $11.75/hr |
| GLM-5.2 | 753.3B | 1684 GB | MIT | No live fit | – |
| MiniMax-M2.7 | 228.7B | 256 GB | Custom license | RTX 4080 Super × 8 | $3.54/hr |
| DeepSeek-V3-0324 | 684.5B | 765 GB | MIT | RTX PRO 6000 WS × 8 | $11.75/hr |
| DeepSeek-V3 | 684.5B | 765 GB | Not stated | RTX PRO 6000 WS × 8 | $11.75/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. See all models for the rest.
Sources
- https://huggingface.co/mistralai/Mistral-Small-24B-Instruct-2501
- https://huggingface.co/microsoft/Phi-4-mini-instruct
- https://huggingface.co/openai/gpt-oss-120b
- https://huggingface.co/openai/gpt-oss-20b
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for reasoning
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: