Qwen2.5 models

56 Qwen2.5 models on Hugging Face, from 135M to 73.4B parameters. At the precision each one is published in, the smallest needs about 0.6 GB of VRAM (turn-detector, cheapest live fit: V100) and the largest about 164 GB (Qwen2.5-VL-72B-Instruct, cheapest live fit: RTX A5000). The cheapest way to run turn-detector is $0.088/hr.

Qwen2.5 models

ModelParametersPublished asVRAM neededLive GPU fitEst. $/hr
turn-detector135MF320.6 GBV100$0.088/hr
jina-code-embeddings-0.5b494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-0.5B494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-0.5B-Instruct494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-Coder-0.5B494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-Coder-0.5B-Instruct494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-0.5B494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-0.5B-Instruct494MBF161.1 GBRTX 3060$0.110/hr
Qwen2.5-Coder-0.5B-Instruct494MBF161.1 GBRTX 3060$0.110/hr
jina-code-embeddings-1.5b1.5BBF163.5 GBRTX 3060$0.110/hr
qwen-base-invoicev1.01-1.5B1.5BBF163.5 GBRTX 3060$0.110/hr
openhands-lm-1.5b-v0.11.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-1.5B1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-1.5B-Instruct1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-Coder-1.5B1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-Coder-1.5B-Instruct1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-Math-1.5B1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-Math-1.5B-Instruct1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-1.5B-Instruct1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-Coder-1.5B-Instruct1.5BBF163.5 GBRTX 3060$0.110/hr
Qwen2.5-3B3.1BBF166.9 GBRTX 3060$0.110/hr
Qwen2.5-3B-Instruct3.1BBF166.9 GBRTX 3060$0.110/hr
Qwen2.5-Coder-3B3.1BBF166.9 GBRTX 3060$0.110/hr
Qwen2.5-Coder-3B-Instruct3.1BBF166.9 GBRTX 3060$0.110/hr
Qwen2.5-3B-Instruct3.1BBF166.9 GBRTX 3060$0.110/hr
VibeThinker-3B3.1BBF166.9 GBRTX 3060$0.110/hr
Nanonets-OCR-s3.8BBF168.4 GBRTX 3060$0.110/hr
Qwen2.5-VL-3B-Instruct3.8BBF168.4 GBRTX 3060$0.110/hr
LocateAnything-3B3.8BBF168.6 GBRTX 3060$0.110/hr
ruadapt_qwen2.5_7B_ext_u48_instruct7.6BBF1616.9 GBRTX A5000$0.176/hr
DeepHat-V1-7B7.6BBF1617.0 GBRTX A5000$0.176/hr
rank1-7b7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-7B7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-7B-Instruct7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-7B-Instruct-1M7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-Coder-7B7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-Coder-7B-Instruct7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-Math-7B7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-Math-7B-Instruct7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-7B-Instruct7.6BBF1617.0 GBRTX A5000$0.176/hr
Qwen2.5-Coder-7B-Instruct7.6BBF1617.0 GBRTX A5000$0.176/hr
NuMarkdown-8B-Thinking8.3BBF1618.5 GBRTX A5000$0.176/hr
Qwen2.5-VL-7B-Instruct8.3BBF1618.5 GBRTX A5000$0.176/hr
Qwen2.5-14B14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-14B-Instruct14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-14B-Instruct-1M14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-Coder-14B14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-Coder-14B-Instruct14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-14B-Instruct14.8BBF1633.0 GBRTX A6000$0.363/hr
Qwen2.5-32B32.8BBF1673.2 GBA100$1.31/hr
Qwen2.5-32B-Instruct32.8BBF1673.2 GBA100$1.31/hr
Qwen2.5-Coder-32B-Instruct32.8BBF1673.2 GBA100$1.31/hr
Qwen2.5-VL-32B-Instruct33.5BBF1674.8 GBA100$1.31/hr
Qwen2.5-72B72.7BBF16163 GBRTX A5000 × 7$1.23/hr
Qwen2.5-72B-Instruct72.7BBF16163 GBRTX A5000 × 7$1.23/hr
Qwen2.5-VL-72B-Instruct73.4BBF16164 GBRTX A5000 × 7$1.23/hr

VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

More from Qwen

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.