What GPU do I need to run it?

VRAM requirements and the cheapest live GPU fit for 228 Hugging Face models, at native, FP8 and INT4 precision. Pick a model below, or jump straight to an owner.

How the VRAM and cost numbers are calculated

  • Weight bytes = parameter count × bytes-per-parameter for the shown precision (FP32=4, BF16/FP16=2, FP8=1, INT4=0.5 bytes).
  • Required VRAM adds a flat 1.2× overhead on top of weight size for KV-cache, activations, and allocator fragmentation, calibrated for an assumed 8,192-token context — a rule-of-thumb, not a per-request simulation. Long contexts or large batch sizes need materially more headroom.
  • “Cheapest live fit” only matches a GPU to a row if its hardware actually supports that row's compute format — a pre-Ampere card (Pascal/Volta/Turing) never qualifies for a BF16 or FP8 row, regardless of VRAM or price — and caps a fit at 8 GPUs of the same model. See the GPU Index for per-model specs.
  • The INT4 row additionally requires a GPU with AWQ/GPTQ/ Marlin-class quantized-serving kernel support AND a quantized checkpoint that's actually been published for the model in question — neither is guaranteed, so check the model's Hugging Face page for a published INT4/AWQ/GPTQ variant before relying on that row.
  • Prices are live per-GPU $/hr rates from Aquanode's marketplace feed, normalized the same way as the GPU Index, with outlier offers excluded.
  • Multi-GPU fits assume the model shards cleanly across cards (e.g. tensor parallelism) — real throughput and interconnect overhead aren't modeled.

Qwen59 models

Qwen3-0.6B
752M params · BF16
Qwen3-8B
8.2B params · BF16
Qwen2.5-1.5B-Instruct
1.5B params · BF16
Qwen2.5-7B-Instruct
7.6B params · BF16
Qwen3-Embedding-0.6B
596M params · BF16
Qwen3-32B
32.8B params · BF16
Qwen3-1.7B
2.0B params · BF16
Qwen2.5-3B-Instruct
3.1B params · BF16
Qwen2.5-0.5B-Instruct
494M params · BF16
Qwen3-4B
4.0B params · BF16
Qwen3-Embedding-8B
7.6B params · BF16
Qwen2.5-14B-Instruct
14.8B params · BF16
Qwen3-4B-Instruct-2507
4.0B params · BF16
Qwen2-1.5B-Instruct
1.5B params · BF16
Qwen3-30B-A3B
30.5B params · BF16
Qwen3-Embedding-4B
4.0B params · BF16
Qwen2.5-Coder-14B-Instruct
14.8B params · BF16
Qwen3-14B
14.8B params · BF16
Qwen3-Reranker-0.6B
596M params · BF16
Qwen3-Reranker-4B
4.0B params · BF16
Qwen2.5-0.5B
494M params · BF16
Qwen3-Coder-Next-FP8
79.7B params · F8_E4M3
Qwen2.5-Coder-7B-Instruct
7.6B params · BF16
Qwen3-30B-A3B-Instruct-2507
30.5B params · BF16
Qwen2.5-32B-Instruct
32.8B params · BF16
Qwen-72B
72.3B params · BF16
Qwen3-0.6B-FP8
752M params · F8_E4M3
Qwen3-Coder-30B-A3B-Instruct
30.5B params · BF16
Qwen3-Coder-30B-A3B-Instruct-FP8
30.5B params · F8_E4M3
Qwen2.5-Coder-32B-Instruct
32.8B params · BF16
Qwen2-0.5B
494M params · BF16
Qwen3-4B-Instruct-2507-FP8
4.4B params · F8_E4M3
Qwen2.5-Coder-7B
7.6B params · BF16
Qwen2.5-7B
7.6B params · BF16
Qwen3-8B-FP8
8.2B params · F8_E4M3
Qwen3-1.7B-Base
1.7B params · BF16
Qwen2.5-1.5B
1.5B params · BF16
Qwen3-0.6B-Base
596M params · BF16
Qwen3-235B-A22B
235.1B params · BF16
Qwen3-4B-Base
4.0B params · BF16
Qwen2.5-Coder-1.5B-Instruct
1.5B params · BF16
Qwen3-Coder-Next
79.7B params · BF16
Qwen2-7B-Instruct
7.6B params · BF16
Qwen3-Reranker-8B
8.2B params · BF16
Qwen3-30B-A3B-Instruct-2507-FP8
30.5B params · F8_E4M3
Qwen2.5-72B-Instruct
72.7B params · BF16
Qwen3-14B-FP8
14.8B params · F8_E4M3
Qwen3-8B-Base
8.2B params · BF16
QwQ-32B
32.8B params · BF16
Qwen3-Next-80B-A3B-Instruct-FP8
81.3B params · F8_E4M3
Qwen3-4B-Thinking-2507
4.0B params · BF16
Qwen2-0.5B-Instruct
494M params · BF16
Qwen3-Coder-480B-A35B-Instruct-FP8
480.2B params · F8_E4M3
Qwen2.5-3B
3.1B params · BF16
Qwen3Guard-Gen-4B
4.4B params · BF16
Qwen3-Next-80B-A3B-Instruct
81.3B params · BF16
Qwen2.5-Coder-1.5B
1.5B params · BF16
Qwen3-4B-Thinking-2507-FP8
4.4B params · F8_E4M3
Qwen1.5-MoE-A2.7B
14.3B params · BF16

meta-llama14 models

deepseek-ai19 models

openai-community3 models

google9 models

zai-org6 models

nvidia12 models

microsoft9 models

HuggingFaceTB5 models

ibm-granite4 models

dphn2 models

EleutherAI4 models

RedHatAI6 models

vikhyatk1 model

h2oai2 models

unsloth7 models

farbodtavakkoli1 model

mistralai3 models

deepreinforce-ai3 models

ibm-research2 models

MiniMaxAI2 models

bigscience2 models

NousResearch4 models

apple1 model

allenai3 models

XiaomiMiMo3 models

tiiuae1 model

Alibaba-NLP1 model

state-spaces1 model

Bahushruth1 model

openbmb1 model

livekit1 model

dots-studio2 models

TIGER-Lab1 model

LiquidAI1 model

shibing6241 model

speakleash1 model

LGAI-EXAONE1 model

typhoon-ai1 model

mlabonne1 model

Zyphra1 model

IlyaGusev1 model

GSAI-ML1 model

Vikhrmodels1 model

rinna1 model

bosonai1 model

lightonai1 model

hmellor1 model

ATH-MaaS1 model

QCRI1 model

ai4bharat1 model

swiss-ai1 model

huggyllama1 model

moondream1 model

google-t51 model

droplychee1 model

JackFram1 model

kenpath1 model

AEON-71 model

codellama1 model

yuxinlu11 model

z-lab1 model

bineric1 model

OpenMOSS-Team1 model

lmms-lab1 model

t-tech1 model

bigcode1 model

Ready when you are

Stop paying for
idle GPUs.

Sign up in 60 seconds. Pay only for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.