Open-weight AI models you can run on a GPU pod
VRAM requirements and the cheapest live GPU fit for 1,429 Hugging Face models, grouped by model line and publisher, at native, FP8 and INT4 precision. Aquanode doesn't run a hosted inference API: every model page links to a GPU pod sized for it, billed per second, that you install your own inference stack on (or, for image and video models, our ComfyUI template). Not sure which model to run in the first place? See how the leading open-weight families score on published benchmarks in the LLM benchmarks leaderboard.
How the VRAM and cost numbers are calculated
- Weight bytes = parameter count × bytes-per-parameter for the shown precision (FP32=4, BF16/FP16=2, FP8=1, INT4=0.5 bytes).
- Required VRAM adds a flat 1.2× overhead on top of weight size for activations and allocator fragmentation, plus KV-cache for text and text+image (VLM) models, calibrated for an assumed 8,192-token context. It's a rule-of-thumb, not a per-request simulation: long contexts or large batch sizes need materially more headroom for language models, diffusion/video models scale mainly with output resolution and frame count instead of context length, and speech models scale mainly with input audio length.
- “Cheapest live fit” only matches a GPU to a row if its hardware actually supports that row's compute format, and caps a fit at 8 GPUs of the same model. A pre-Ampere card (Pascal/Volta/Turing) never qualifies for a BF16 or FP8 row, regardless of VRAM or price. See the GPU Index for per-model specs.
- The INT4 row additionally requires a quantized checkpoint that's actually been published for the model in question. For language models this also means a GPU with AWQ/GPTQ/Marlin-class quantized-serving kernel support. Neither is guaranteed, so check the model's Hugging Face page for a published quantized variant before relying on that row.
- Prices are live per-GPU $/hr rates from Aquanode's live price feed, normalized the same way as the GPU Index, with outlier offers excluded.
- Multi-GPU fits assume the model shards cleanly across cards (e.g. tensor parallelism). Real throughput and interconnect overhead aren't modeled.
Qwen
Meta
Stability AI
NVIDIA
Z.ai
DeepSeek
Meta
IBM
Microsoft
OpenAI
Ai2
EleutherAI
Hugging Face
ornith-ai
Wan-AI
Liquid AI
Mistral AI
TII
Moonshot AI
OpenBMB
Tencent
01.AI
MiniMax
Black Forest Labs
inclusionAI
llava-hf
openai-community
LG AI Research
OpenGVLab
internlm
Nous Research
ATH-MaaS
bigscience
Cohere
typhoon-ai
upstage
Xiaomi MiMo
Lightricks
Salesforce
arcee-ai
Baidu
KBlueLeaf
parler-tts
poolside
StepFun
Zyphra
ai21labs
ByteDance
ByteDance Seed
convaiinnovations
datalab-to
fastino
FastVideo
google-t5
h2oai
HuggingFaceM4
lightonai
meituan-longcat
Nanbeige
skt
state-spaces
swiss-ai
utter-project
vikhyatk
More models
Single models and small groups of models from one publisher.
- bert-base-uncased google-bert
- MiniMaxAI MiniMax (2 models)
- DeepSeek OCR DeepSeek (2 models)
- Unlimited-OCR Baidu
- zai-org Z.ai (2 models)
- Voxtral-Mini-4B-Realtime-2602 Mistral AI
- phi-2 Microsoft
- multilingual-e5-large-instruct intfloat
- OpenELM-1_1B-Instruct apple
- google Google (2 models)
- GPT-NeoX EleutherAI (2 models)
- ibm-research ibm-research (2 models)
- Lykon Lykon (2 models)
- Rax-4.5 raxcore-dev
- dots-studio dots-studio (2 models)
- Tongyi-MAI Tongyi-MAI (2 models)
- tencent Tencent (2 models)
- distil-whisper distil-whisper (2 models)
- Muse-Glimmer-30B meta-models
- macbert4csc-base-chinese shibing624
- BAAI BAAI (2 models)
- olmOCR Ai2 (2 models)
- Infinity-Parser2-Pro infly
- saiga_llama3_8b IlyaGusev
- opendatalab opendatalab (2 models)
- thinkingmachines thinkingmachines (2 models)
- openbmb OpenBMB (2 models)
- RealVisXL_V5.0 SG161222
- InternScience InternScience (2 models)
- nvidia NVIDIA (2 models)
- vllm-translategemma-4b-it Infomaniak-AI
- playground-v2.5-1024px-aesthetic playgroundai
- ChatRex-7B IDEA-Research
- Kimi-VL-A3B-Instruct Moonshot AI
- s2-pro fishaudio
- Llama 4 Meta (2 models)
- Fanar-1-9B-Instruct QCRI
- voyage-4-nano voyageai
- Stable Video Diffusion Stability AI (2 models)
- MOSS-Transcribe-Diarize OpenMOSS-Team
- Mixtral Mistral AI (2 models)
- bosonai bosonai (2 models)
- openai-gpt openai-community
- Kimi-Linear-48B-A3B-Instruct Moonshot AI
- LLaMmlein_1B_prerelease LSX-UniWue
- vllm-translategemma-12b-it chbae624
- StarCoder2 bigcode (2 models)
- allenai Ai2 (2 models)
- Ilama-3.2-1B hmellor
- Qari-OCR-v0.3-VL-2B-Instruct NAMAA-Space
- krea krea (2 models)
- XCurOS XCurOS (2 models)
- pixtral-12b mistral-experimental
- mixedbread-ai mixedbread-ai (2 models)
- lodestones lodestones (2 models)
- LCM_Dreamshaper_v7 SimianLuo
- zeta-2.1-autoround-W4A16 LeaderboardModel1
- llm-jp llm-jp (2 models)
- Dream-org Dream-org (2 models)
- MiMo Xiaomi MiMo (2 models)
- openvla-7b openvla
- hf-moshiko kmhf
- transformer-1.3B-100B fla-hub
- XingChen-AGI XingChen-AGI (2 models)
- OneRec-1.7B OpenOneRec
- HiDream-ai HiDream-ai (2 models)
- internlm3-8b-instruct internlm
- sarashina2.2-0.5b-instruct-v0.1 sbintuitions
- Phi-1 Microsoft (2 models)
- openthaigpt1.5-7b-instruct openthaigpt
- Moonlight Moonshot AI (2 models)
- QwQ Qwen (2 models)
- Falcon Mamba TII (2 models)
- Bielik-11B-v3.0-Instruct speakleash
- OLMo Ai2 (2 models)
- Hunyuan Tencent (2 models)
- Phi-mini-MoE-instruct Microsoft
- progen2-small hugohrban
- Vikhr-Nemo-12B-Instruct-R-21-09-24 Vikhrmodels
- ideogram-4-fp8 ideogram-ai
- DeepSeek MoE DeepSeek (2 models)
- digiplay digiplay (2 models)
- Tongyi-DeepResearch-30B-A3B Alibaba-NLP
- CMSManhattan CMSManhattan (2 models)
- Command R Cohere (2 models)
- KAT-Coder-V2.5-Dev Kwaipilot
- sundial-base-128m thuml
- lumeleto gratefulasi
- K-intelligence K-intelligence (2 models)
- gpt_bigcode-santacoder bigcode
- Audio8-ASR-Infinite Edge0
- briaai briaai (2 models)
- Salesforce Salesforce (2 models)
- Lite-Oute-1-300M OuteAI
- Canary NVIDIA (2 models)
- stella_en_1.5B_v5 NovaSearch
- YuE2-3B m-a-p
- sensenova sensenova (2 models)
- paloalma paloalma (2 models)
- UnfilteredAI UnfilteredAI (2 models)
- lucid-v1-nemo dreamgen
- HRM-Text-1B sapientinc
- JoyAI-LLM-Flash-MXFP8-last-6-BF16 zianglih
- ChemLLM-7B-Chat-1_5-DPO AI4Chem
- Dia-1.6B nari-labs
- kyutai kyutai (2 models)
- CrisperWhisper2.0_large nyralabs
- Janus-Pro-1B deepseek-community
- SDAR-1.7B-Chat JetLM
- salamandra-7b-instruct BSC-LT
- syvai syvai (2 models)
- NeuralMonarch-7B mlabonne
- functiongemma-270m-it Google
- Audio8 Audio8 (2 models)
- Fun-ASR-Nano-2512-hf FunAudioLLM
- kakaocorp kakaocorp (2 models)
- LUSTIFY-v2.0 shootstuff
- K-EXAONE-236B-A23B LG AI Research
- IFM IFM (2 models)
- Saul-7B-Instruct-v1 Equall
- text-to-video-ms-1.7b ali-vilab
- Intern-S1 internlm
- CheXagent-2-3b StanfordAIMI
- A.X-K2-Raon-Speech-21B-A3B KRAFTON
- Mistral-Nemo-Instruct-2407-lenient-chatfix m8than
- huginn-0125 tomg-group-umd
- Ministral-3b-instruct ministral
- fg-clip-base qihoo360
- KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS KridgeDookie
- LUSTIFY-v2.0 TheImposterImposters
- zeta-2.1 zed-industries
- MiniMax-Text-01-hf MiniMax
- Mellum2-12B-A2.5B-Base JetBrains
- MixTAO-7Bx2-MoE-v8.1 mixtao
- rho-1b-sft-MATH realtreetune
- aya-expanse-8b Cohere
- NSFW-Uncensored Heartsync
- SSD-1B segmind
- Juggernaut-Z-Image RunDiffusion
- BanglaASR bangla-speech-processing
- babylm-multimodal-baseline-flamingo BabyLM-community
- GritLM-7B-vllm parasail-ai
- AV-HuBERT-MuAViC-en nguyenvulebinh
- pornmasterPro_noobV3VAE votepurchase
- sarvam-30b sarvamai
- Bamba-9B-v1 ibm-ai-platform
- pixai-tagger-v1.0 pixai-labs
- sqlcoder-7b-2 defog
- Veena maya-research
- MUSE-news_target muse-bench
- PixelModel-v6 bench-labs
- GigaChat3-10B-A1.8B ai-sage
- ZDTaichu5.0-9B TaichuAI
- jetmoe-8b jetmoe
- Param2-17B-A2.4B-Thinking bharatgenai
- MagicPrompt-Stable-Diffusion Gustavosta
- Counterfeit-V2.5 gsdf
- AFM-4.5B arcee-ai
- lingbot-world-fast robbyant
- ThinkPRM-1.5B launch
- prometheus-7b-v2.0 prometheus-eval
- acestep-5Hz-lm-4B ACE-Step
- mochi-1-preview genmo
- DFM-Mimir danish-foundation-models
- muscriptor-large MuScriptor
- Hy4 Tencent (2 models)
- openjev openjev
- Aleph-Alpha Aleph-Alpha (2 models)
- AliceAI-Foundation-80B-A3B-Base yandex
- Julia-1 SupersonicLabs
- Breeze-TTS-2 BreezeBlue
- IndexTeam IndexTeam (2 models)
- ReaderLM-v2 jinaai
- IQuest-Q1 IQuestLab
- LeVJEPA-VideoMix-Large galilai-group
- Darwin-180B-RSI FINAL-Bench
- XHToken XHToken (2 models)
- ASL-4B-v1 saai-sa
- perplexity-ai perplexity-ai (2 models)
- FRIDA-Decisions ai-forever
- sopro-v2-turbo samuel-vitorino
- baguettotron-600m PleIAs
- kandinskylab kandinskylab (2 models)
- gnani-evon-v3.3-30B-A3B gnani
- ice-012-audio darkps
- Isaac-0.5 PerceptronAI
- Wayfarer-12B LatitudeGames
- Irodori-TTS-v4-Large Aratako
- DeepSeek V4 DeepSeek (2 models)
- Ming-Image-0.1-Design inclusionAI
- timesfm-3.0-pytorch Google