Best open reasoning models
How to pick an open reasoning model: thinking output length, license, size and context. Picks by GPU budget, with facts from model cards.
A reasoning model writes a long chain of thought before it answers, which helps on math, code, and multi-step problems and costs you extra generated tokens on every request. Choosing one is mostly about how much thinking you can afford, the license, and the size that fits your GPU. The catalog below lists the rest.
What matters when picking a reasoning model
Thinking tokens are the real cost. The visible answer may be short, but the reasoning before it can run to thousands of tokens. DeepSeek's cards set a 64K maximum generation length for their evaluations, which shows how long these traces can get. Plan for slow, memory-hungry generation: the KV cache fills as the trace grows.
Distilled or original. The full DeepSeek-R1 is a 671B mixture-of-experts model with 37B activated parameters and a 128K context, per its card. DeepSeek also released smaller dense models fine-tuned on reasoning data generated by R1. That is distillation, and it puts reasoning behavior into sizes a single GPU can hold.
Reasoning effort controls. Some models let you set how long they think, which lets one model serve both quick and hard requests.
License. The DeepSeek-R1 cards say the series supports commercial use and distillation under the MIT License. Check each card, since the distilled models are built on other base models.
Precision. BF16 is the reference. A quantized build cuts the footprint at some accuracy cost. The VRAM calculator helps size it.
Picks by situation
Smaller GPUs (a 24 GB card): DeepSeek-R1-0528-Qwen3-8B. DeepSeek distilled the chain of thought from DeepSeek-R1-0528 into Qwen3 8B Base and reports it surpasses Qwen3 8B by 10.0% on AIME 2024 (vendor-reported). It is MIT licensed and has about 8.2B parameters in the catalog.
Mid-size dense: DeepSeek-R1-Distill-Qwen-32B. Built on Qwen2.5-32B and MIT licensed. DeepSeek says it outperforms OpenAI-o1-mini across various benchmarks (vendor-reported). Phi-4-reasoning is a smaller alternative: Phi-4-reasoning is a 14B dense model with a 32k context under the MIT license, according to Microsoft's card.
One 80 GB GPU: gpt-oss-120b. OpenAI's card lists 117B parameters with 5.1B active, Apache 2.0, and says it runs on a single 80GB GPU thanks to MXFP4 quantization of the mixture-of-experts weights. It has configurable reasoning effort (low, medium, high) and exposes the full chain of thought. Its smaller sibling gpt-oss-20b has 21B parameters with 3.6B active and targets lower latency.
Largest dense distill: DeepSeek-R1-Distill-Llama-70B. A 70B distillation under the MIT license on the card. It needs more than one consumer-class card in BF16, so plan for a multi-GPU rental or a quantized build.
Most permissive license: the DeepSeek-R1 family and Phi-4-reasoning are MIT. The gpt-oss models are Apache 2.0.
Serving notes
Reasoning workloads favor GPUs with high memory bandwidth, because every thinking token is a decode step. Stream the output so users are not staring at a blank screen, and cap the thinking budget where your application allows it.
Aquanode rents GPUs by the hour, so you can match the card to the model and stop when the job is done. See pricing. Benchmark numbers above come from the vendors' own cards; run your own prompts before choosing.
Open models for reasoning
All 29 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| MiniCPM-V-4.6-Thinking | 1.3B | 2.9 GB | – | RTX 4070 Super | $0.121/hr |
| Ouro-2.6B-Thinking | 2.7B | 6.0 GB | – | RTX 4070 Super | $0.121/hr |
| Ouro-1.4B-Thinking | 1.4B | 3.2 GB | – | RTX 4070 Super | $0.121/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| DeepSeek-R1-0528-Qwen3-8B | 8.2B | 18.3 GB | – | RTX A5000 | $0.176/hr |
| Qwen3-4B-Thinking-2507 | 4.0B | 9.0 GB | – | RTX 4070 Super | $0.121/hr |
| Qwen3-VL-8B-Thinking | 8.8B | 19.6 GB | – | RTX A5000 | $0.176/hr |
| Olmo-3-7B-Think | 7.3B | 16.3 GB | – | RTX A5000 | $0.176/hr |
| Phi-4-mini-reasoning | 3.8B | 8.6 GB | – | RTX 4070 Super | $0.121/hr |
| Nemotron-H-8B-Reasoning-128K | 8.1B | 18.1 GB | – | RTX A5000 | $0.176/hr |
| AI21-Jamba-Reasoning-3B | 3.2B | 7.1 GB | – | RTX 4070 Super | $0.121/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| GLM-4.1V-9B-Thinking | 10.3B | 23.0 GB | – | RTX A5000 | $0.176/hr |
| Qwen3-30B-A3B-Thinking-2507 | 30.5B | 68.2 GB | – | A100 | $1.21/hr |
| QwQ-32B | 32.8B | 73.2 GB | – | A100 | $1.21/hr |
| Nemotron-Cascade-2-30B-A3B | 31.6B | 70.6 GB | – | A100 | $1.21/hr |
| HyperCLOVAX-SEED-Think-32B | 33.3B | 74.5 GB | – | A100 | $1.21/hr |
| HyperCLOVAX-SEED-Think-14B | 14.7B | 65.9 GB | – | A100 | $1.21/hr |
| Phi-4-reasoning | 14.7B | 32.8 GB | – | RTX A6000 | $0.363/hr |
| GLM-Z1-32B-0414 | 32.6B | 72.8 GB | – | A100 | $1.21/hr |
| QwQ-32B-Preview | 32.8B | 73.2 GB | – | A100 | $1.21/hr |
| ERNIE-4.5-21B-A3B-Thinking | 21.8B | 48.8 GB | – | A100 | $1.21/hr |
| llm-jp-4-33b-thinking | 33.2B | 74.3 GB | – | A100 | $1.21/hr |
| Param2-17B-A2.4B-Thinking | 17.2B | 38.3 GB | – | RTX A6000 | $0.363/hr |
| Phi-4-reasoning-plus | 14.7B | 32.8 GB | – | RTX A6000 | $0.363/hr |
| gpt-oss-20b | 20.9B | 16.0 GB | Apache 2.0 | RTX A5000 | $0.176/hr |
40B to 150B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen3-Next-80B-A3B-Thinking | 81.3B | 182 GB | Apache 2.0 | RTX A5000 × 8 | $1.41/hr |
| gpt-oss-120b | 116.8B | 80.0 GB | Apache 2.0 | A100 | $1.21/hr |
150B parameters and up
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| DeepSeek-R1 | 684.5B | 765 GB | – | RTX PRO 6000 × 8 | $11.75/hr |
| DeepSeek-R1-0528 | 684.5B | 765 GB | MIT | RTX PRO 6000 × 8 | $11.75/hr |
| Qwen3-235B-A22B-Thinking-2507 | 235.1B | 525 GB | Apache 2.0 | RTX PRO 6000 × 6 | $8.25/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
- https://huggingface.co/deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
- https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
- https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B
- https://huggingface.co/microsoft/Phi-4-reasoning
- https://huggingface.co/openai/gpt-oss-120b
- https://huggingface.co/openai/gpt-oss-20b
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: