Best open reasoning models

How to pick an open reasoning model: thinking output length, license, size and context. Picks by GPU budget, with facts from model cards.

A reasoning model writes a long chain of thought before it answers, which helps on math, code, and multi-step problems and costs you extra generated tokens on every request. Choosing one is mostly about how much thinking you can afford, the license, and the size that fits your GPU. The catalog below lists the rest.

What matters when picking a reasoning model

Thinking tokens are the real cost. The visible answer may be short, but the reasoning before it can run to thousands of tokens. DeepSeek's cards set a 64K maximum generation length for their evaluations, which shows how long these traces can get. Plan for slow, memory-hungry generation: the KV cache fills as the trace grows.

Distilled or original. The full DeepSeek-R1 is a 671B mixture-of-experts model with 37B activated parameters and a 128K context, per its card. DeepSeek also released smaller dense models fine-tuned on reasoning data generated by R1. That is distillation, and it puts reasoning behavior into sizes a single GPU can hold.

Reasoning effort controls. Some models let you set how long they think, which lets one model serve both quick and hard requests.

License. The DeepSeek-R1 cards say the series supports commercial use and distillation under the MIT License. Check each card, since the distilled models are built on other base models.

Precision. BF16 is the reference. A quantized build cuts the footprint at some accuracy cost. The VRAM calculator helps size it.

Picks by situation

Smaller GPUs (a 24 GB card): DeepSeek-R1-0528-Qwen3-8B. DeepSeek distilled the chain of thought from DeepSeek-R1-0528 into Qwen3 8B Base and reports it surpasses Qwen3 8B by 10.0% on AIME 2024 (vendor-reported). It is MIT licensed and has about 8.2B parameters in the catalog.

Mid-size dense: DeepSeek-R1-Distill-Qwen-32B. Built on Qwen2.5-32B and MIT licensed. DeepSeek says it outperforms OpenAI-o1-mini across various benchmarks (vendor-reported). Phi-4-reasoning is a smaller alternative: Phi-4-reasoning is a 14B dense model with a 32k context under the MIT license, according to Microsoft's card.

One 80 GB GPU: gpt-oss-120b. OpenAI's card lists 117B parameters with 5.1B active, Apache 2.0, and says it runs on a single 80GB GPU thanks to MXFP4 quantization of the mixture-of-experts weights. It has configurable reasoning effort (low, medium, high) and exposes the full chain of thought. Its smaller sibling gpt-oss-20b has 21B parameters with 3.6B active and targets lower latency.

Largest dense distill: DeepSeek-R1-Distill-Llama-70B. A 70B distillation under the MIT license on the card. It needs more than one consumer-class card in BF16, so plan for a multi-GPU rental or a quantized build.

Most permissive license: the DeepSeek-R1 family and Phi-4-reasoning are MIT. The gpt-oss models are Apache 2.0.

Serving notes

Reasoning workloads favor GPUs with high memory bandwidth, because every thinking token is a decode step. Stream the output so users are not staring at a blank screen, and cap the thinking budget where your application allows it.

Aquanode rents GPUs by the hour, so you can match the card to the model and stop when the job is done. See pricing. Benchmark numbers above come from the vendors' own cards; run your own prompts before choosing.

Open models for reasoning

All 29 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
MiniCPM-V-4.6-Thinking1.3B2.9 GB–RTX 4070 Super$0.121/hr
Ouro-2.6B-Thinking2.7B6.0 GB–RTX 4070 Super$0.121/hr
Ouro-1.4B-Thinking1.4B3.2 GB–RTX 4070 Super$0.121/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
DeepSeek-R1-0528-Qwen3-8B8.2B18.3 GB–RTX A5000$0.176/hr
Qwen3-4B-Thinking-25074.0B9.0 GB–RTX 4070 Super$0.121/hr
Qwen3-VL-8B-Thinking8.8B19.6 GB–RTX A5000$0.176/hr
Olmo-3-7B-Think7.3B16.3 GB–RTX A5000$0.176/hr
Phi-4-mini-reasoning3.8B8.6 GB–RTX 4070 Super$0.121/hr
Nemotron-H-8B-Reasoning-128K8.1B18.1 GB–RTX A5000$0.176/hr
AI21-Jamba-Reasoning-3B3.2B7.1 GB–RTX 4070 Super$0.121/hr

10B to 40B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
GLM-4.1V-9B-Thinking10.3B23.0 GB–RTX A5000$0.176/hr
Qwen3-30B-A3B-Thinking-250730.5B68.2 GB–A100$1.21/hr
QwQ-32B32.8B73.2 GB–A100$1.21/hr
Nemotron-Cascade-2-30B-A3B31.6B70.6 GB–A100$1.21/hr
HyperCLOVAX-SEED-Think-32B33.3B74.5 GB–A100$1.21/hr
HyperCLOVAX-SEED-Think-14B14.7B65.9 GB–A100$1.21/hr
Phi-4-reasoning14.7B32.8 GB–RTX A6000$0.363/hr
GLM-Z1-32B-041432.6B72.8 GB–A100$1.21/hr
QwQ-32B-Preview32.8B73.2 GB–A100$1.21/hr
ERNIE-4.5-21B-A3B-Thinking21.8B48.8 GB–A100$1.21/hr
llm-jp-4-33b-thinking33.2B74.3 GB–A100$1.21/hr
Param2-17B-A2.4B-Thinking17.2B38.3 GB–RTX A6000$0.363/hr
Phi-4-reasoning-plus14.7B32.8 GB–RTX A6000$0.363/hr
gpt-oss-20b20.9B16.0 GBApache 2.0RTX A5000$0.176/hr

40B to 150B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-Next-80B-A3B-Thinking81.3B182 GBApache 2.0RTX A5000 × 8$1.41/hr
gpt-oss-120b116.8B80.0 GBApache 2.0A100$1.21/hr

150B parameters and up

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
DeepSeek-R1684.5B765 GB–RTX PRO 6000 × 8$11.75/hr
DeepSeek-R1-0528684.5B765 GBMITRTX PRO 6000 × 8$11.75/hr
Qwen3-235B-A22B-Thinking-2507235.1B525 GBApache 2.0RTX PRO 6000 × 6$8.25/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.