Best open chat models

How to pick an open chat model: license, context, size and latency. Picks by GPU budget with facts taken from the publishers' model cards.

For a chat assistant you want an instruct-tuned model that follows instructions, holds a conversation, and answers fast. Pick on license, context length, and size, then test it on your own prompts. This page covers how to choose; the catalog below lists the instruct models.

What matters when picking a chat model

Instruct, not base. Base checkpoints continue text and do not reliably follow instructions. Chat apps want the instruct (or "it", "chat") variant from the same publisher.

Latency and concurrency. Chat is interactive, so time to first token and decode speed shape the experience. Smaller models, and mixture-of-experts models that run only some weights per token, respond faster. If many users share one GPU, the KV cache for each conversation also eats memory.

Context length. Long chats and document questions need room. Longer context costs memory, so match it to what your product really sends.

Tool calling and structured output. If the assistant calls functions or returns JSON, check that the card advertises native support.

License. Apache 2.0 and MIT cards carry no extra use restrictions of their own. Some chat models use custom community licenses with conditions, so read the card.

Precision. BF16 is the reference; a quantized build shrinks the footprint. Use the VRAM calculator to size it.

Picks by situation

Smallest and cheapest to serve: Phi-4-mini-instruct. Microsoft's card lists a 128K context and an MIT license, and the catalog shows about 3.8B parameters. It is a good fit for high-volume, low-latency assistants where a small model is enough.

Smaller GPUs (a 24 GB card): Mistral-Small-24B-Instruct-2501. The card lists 24B parameters, a 32k context window, and Apache 2.0, and highlights native function calling and JSON output. A 24B model in BF16 is a stretch for 24 GB, so expect to run a quantized build there or step up a card size.

One 80 GB GPU: gpt-oss-120b. OpenAI's card lists 117B parameters with 5.1B active and Apache 2.0, and says it fits a single 80GB GPU. Only a fraction of the weights run per token, which helps latency. It also has configurable reasoning effort, so you can keep chat replies quick or let it think longer on hard questions.

Lower latency from the same family: gpt-oss-20b. 21B parameters with 3.6B active, Apache 2.0, aimed at lower-latency and local use per the card.

Longest context: Phi-4-mini-instruct at 128K is the longest among these picks. For multimodal chat with even longer windows, see the vision-language page.

Most permissive license: Mistral Small and the gpt-oss models are Apache 2.0, and Phi-4-mini-instruct is MIT, all per their cards.

Serving notes

Run the model behind an inference server that batches requests, since chat traffic is bursty. Keep the system prompt short, because it is re-read on every turn. For a first deployment, pick the smallest model that passes your own evaluation set, then scale up only if it fails.

Aquanode rents GPUs by the hour, so you can try a 4B model and a 100B-class model on the right card sizes and keep the one that works. See pricing. Model-card claims are vendor-reported, so measure on your own traffic.

Open models for chat and assistants

The 60 most downloaded of 473 models in the catalog, grouped by size. The rest are listed on the models directory.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-0.6B752M1.7 GB–RTX 4070 Super$0.121/hr
gpt2137M0.6 GB–V100$0.088/hr
Qwen2.5-1.5B-Instruct1.5B3.5 GB–RTX 4070 Super$0.121/hr
Llama-3.2-1B-Instruct1.2B2.8 GB–RTX 4070 Super$0.121/hr
Qwen2.5-0.5B-Instruct494M1.1 GB–RTX 4070 Super$0.121/hr
Qwen3-1.7B2.0B4.5 GBApache 2.0RTX 4070 Super$0.121/hr
pythia-160m213M0.5 GB–V100$0.088/hr
gemma-3-1b-it1000M2.2 GB–RTX 4070 Super$0.121/hr
SmolLM2-135M135M0.3 GB–RTX 4070 Super$0.121/hr
Qwen2.5-0.5B494M1.1 GB–RTX 4070 Super$0.121/hr
phi-22.8B6.2 GB–V100$0.088/hr
SmolLM2-135M-Instruct135M0.3 GB–RTX 4070 Super$0.121/hr
gemma-3-270m268M0.6 GB–RTX 4070 Super$0.121/hr
OpenELM-1_1B-Instruct1.1B2.4 GB–RTX 4070 Super$0.121/hr
Llama-3.2-1B1.2B2.8 GB–RTX 4070 Super$0.121/hr
Qwen2.5-1.5B1.5B3.5 GB–RTX 4070 Super$0.121/hr
gpt2-large812M3.6 GB–V100$0.088/hr
OLMo-2-0425-1B1.5B6.6 GB–V100$0.088/hr
bloomz-560m559M1.2 GB–V100$0.088/hr
Qwen2-1.5B-Instruct1.5B3.5 GB–RTX 4070 Super$0.121/hr
Qwen2-0.5B494M1.1 GB–RTX 4070 Super$0.121/hr
Qwen2.5-Math-1.5B1.5B3.5 GB–RTX 4070 Super$0.121/hr
MiniCPM5-1B1.1B2.4 GB–RTX 4070 Super$0.121/hr
gemma-2-2b-it2.6B5.8 GB–RTX 4070 Super$0.121/hr
pythia-160m-deduped213M0.5 GB–V100$0.088/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-8B8.2B18.3 GBApache 2.0RTX A5000$0.176/hr
Qwen2.5-7B-Instruct7.6B17.0 GBApache 2.0RTX A5000$0.176/hr
Qwen2.5-3B-Instruct3.1B6.9 GBCustom licenseRTX 4070 Super$0.121/hr
Qwen3-4B4.0B9.0 GBApache 2.0RTX 4070 Super$0.121/hr
Llama-3.1-8B-Instruct8.0B17.9 GBLlama 3.1 Community License AgreementRTX A5000$0.176/hr
Qwen3-4B-Instruct-25074.0B9.0 GB–RTX 4070 Super$0.121/hr
Meta-Llama-3-8B-Instruct8.0B17.9 GB–RTX A5000$0.176/hr
Llama-3.2-3B-Instruct3.2B7.2 GBLlama 3.2 Community License AgreementRTX 4070 Super$0.121/hr
Mistral-7B-Instruct-v0.27.2B16.2 GB–RTX A5000$0.176/hr
PowerMoE-3b3.4B15.1 GB–V100$0.088/hr
granite-4.1-8b8.8B19.7 GB–RTX A5000$0.176/hr
DeepSeek-R1-0528-Qwen3-8B8.2B18.3 GB–RTX A5000$0.176/hr
Llama-2-7b-hf6.7B15.1 GB–V100$0.088/hr
gemma-2-9b-it9.2B20.7 GB–RTX A5000$0.176/hr
Meta-Llama-3-8B8.0B17.9 GB–RTX A5000$0.176/hr
Apertus-8B-Instruct-25098.1B18.0 GB–RTX A5000$0.176/hr
Phi-3-mini-4k-instruct3.8B8.5 GBMITRTX 4070 Super$0.121/hr

10B to 40B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-32B32.8B73.2 GBApache 2.0A100$1.21/hr
Qwen2.5-14B-Instruct14.8B33.0 GBApache 2.0RTX A6000$0.363/hr
Qwen3-30B-A3B30.5B68.2 GBApache 2.0A100$1.21/hr
Qwen2.5-32B-Instruct32.8B73.2 GB–A100$1.21/hr
GLM-4.7-Flash31.2B69.8 GBMITA100$1.21/hr
Qwen3-14B14.8B33.0 GB–RTX A6000$0.363/hr
Qwen3-30B-A3B-Instruct-250730.5B68.2 GB–A100$1.21/hr
NVIDIA-Nemotron-3-Nano-30B-A3B-BF1631.6B70.6 GB–A100$1.21/hr
gpt-neox-20b20.7B46.4 GB–RTX A6000$0.363/hr
phi-414.7B32.8 GBMITRTX A6000$0.363/hr

40B to 150B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen-72B72.3B162 GB–RTX A5000 × 7$1.23/hr
NVIDIA-Nemotron-3-Super-120B-A12B-BF16123.6B276 GB–RTX A6000 × 6$2.18/hr

150B parameters and up

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
DeepSeek-R1684.5B765 GB–RTX PRO 6000 WS × 8$11.75/hr
DeepSeek-V3.2685.4B766 GB–RTX PRO 6000 WS × 8$11.75/hr
GLM-5.2753.3B1684 GBMITNo live fit–
MiniMax-M2.7228.7B256 GBCustom licenseRTX 4080 Super × 8$3.54/hr
DeepSeek-V3-0324684.5B765 GBMITRTX PRO 6000 WS × 8$11.75/hr
DeepSeek-V3684.5B765 GBNot statedRTX PRO 6000 WS × 8$11.75/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. See all models for the rest.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.