Open-weight AI models you can run on a GPU pod

VRAM requirements and the cheapest live GPU fit for 1,429 Hugging Face models, grouped by model line and publisher, at native, FP8 and INT4 precision. Aquanode doesn't run a hosted inference API: every model page links to a GPU pod sized for it, billed per second, that you install your own inference stack on (or, for image and video models, our ComfyUI template). Not sure which model to run in the first place? See how the leading open-weight families score on published benchmarks in the LLM benchmarks leaderboard.

How the VRAM and cost numbers are calculated

  • Weight bytes = parameter count × bytes-per-parameter for the shown precision (FP32=4, BF16/FP16=2, FP8=1, INT4=0.5 bytes).
  • Required VRAM adds a flat 1.2× overhead on top of weight size for activations and allocator fragmentation, plus KV-cache for text and text+image (VLM) models, calibrated for an assumed 8,192-token context. It's a rule-of-thumb, not a per-request simulation: long contexts or large batch sizes need materially more headroom for language models, diffusion/video models scale mainly with output resolution and frame count instead of context length, and speech models scale mainly with input audio length.
  • “Cheapest live fit” only matches a GPU to a row if its hardware actually supports that row's compute format, and caps a fit at 8 GPUs of the same model. A pre-Ampere card (Pascal/Volta/Turing) never qualifies for a BF16 or FP8 row, regardless of VRAM or price. See the GPU Index for per-model specs.
  • The INT4 row additionally requires a quantized checkpoint that's actually been published for the model in question. For language models this also means a GPU with AWQ/GPTQ/Marlin-class quantized-serving kernel support. Neither is guaranteed, so check the model's Hugging Face page for a published quantized variant before relying on that row.
  • Prices are live per-GPU $/hr rates from Aquanode's live price feed, normalized the same way as the GPU Index, with outlier offers excluded.
  • Multi-GPU fits assume the model shards cleanly across cards (e.g. tensor parallelism). Real throughput and interconnect overhead aren't modeled.

Qwen

Meta

Google

Stability AI

NVIDIA

Z.ai

DeepSeek

Meta

IBM

Microsoft

OpenAI

Ai2

EleutherAI

Hugging Face

ornith-ai

Wan-AI

Liquid AI

Mistral AI

TII

Moonshot AI

OpenBMB

Tencent

01.AI

MiniMax

Black Forest Labs

inclusionAI

llava-hf

openai-community

LG AI Research

OpenGVLab

internlm

Nous Research

ATH-MaaS

bigscience

Cohere

typhoon-ai

upstage

Xiaomi MiMo

Lightricks

Salesforce

arcee-ai

Baidu

KBlueLeaf

parler-tts

poolside

StepFun

Zyphra

ai21labs

ByteDance

ByteDance Seed

convaiinnovations

datalab-to

fastino

FastVideo

google-t5

h2oai

HuggingFaceM4

lightonai

meituan-longcat

Nanbeige

naver-hyperclovax

skt

state-spaces

swiss-ai

utter-project

vikhyatk

More models

Single models and small groups of models from one publisher.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.