AI models that fit on a 288 GB GPU

Open models that fit in 288 GB of VRAM: 25 as published, 16 at FP8 and 22 at INT4. Each model is listed on the smallest tier it fits at that precision, so anything smaller is on the 192 GB page or below.

Fits as published

Models whose published weights, plus the flat overhead, fit this much memory with no quantization.

Text models

16 models.

ModelParametersPublished asVRAM needed
MiniMax-M2.7228.7BF8_E4M3256 GB
NVIDIA-Nemotron-3-Super-120B-A12B-BF16123.6BBF16276 GB
MiniMax-M2.5228.7BF8_E4M3256 GB
MiniMax-M2228.7BF8_E4M3256 GB
Qwen3-235B-A22B-Instruct-2507-FP8235.1BF8_E4M3263 GB
GLM-4.5-Air110.5BBF16247 GB
Laguna-S-2.1117.6BBF16263 GB
Qwen3-235B-A22B-Thinking-2507-FP8235.1BF8_E4M3263 GB
MiniMax-M2.1228.7BF8_E4M3256 GB
Qwen3-VL-235B-A22B-Instruct-FP8-dynamic235.8BF8_E4M3263 GB
DeepSeek-V2.5-1210-FP8235.7BF8_E4M3263 GB
Ling-3.0-flash127.5BBF16285 GB
Qwen3-235B-A22B-FP8235.1BF8_E4M3263 GB
Solar-Open-100B102.7BBF16229 GB
NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16123.6BBF16276 GB
GLM-4.5-Air-Base110.5BBF16247 GB

Vision-language models

9 models.

ModelParametersPublished asVRAM needed
Qwen3.5-122B-A10B125.1BBF16280 GB
Llama-4-Scout-17B-16E-Instruct108.6BBF16243 GB
Qwen3.8-Flash-Next-FP8180.0BF8_E4M3201 GB
Qwen3-VL-235B-A22B-Instruct-FP8235.7BF8_E4M3263 GB
Llama-3.2-90B-Vision-Instruct88.6BBF16198 GB
GLM-4.5V107.7BBF16241 GB
Qwen3.8-Flash-Next-FP8180.0BF8_E4M3201 GB
Qwen3.8-Flash-Next-Uncensored-FP8180.0BF8_E4M3201 GB
Qwen3.8-Flash-Next-UNCENSORED-FP8180.0BF8_E4M3201 GB

Fits at FP8

Models that fit only after quantizing the weights to 8 bits (1 byte per parameter). Needs an FP8 checkpoint or an engine that quantizes on load, and a GPU with FP8 support.

Text models

10 models.

ModelParametersPublished asVRAM needed
Qwen3-235B-A22B235.1BBF16263 GB
Step-3.5-Flash199.4BBF16223 GB
Qwen3-235B-A22B-Instruct-2507235.1BBF16263 GB
DeepSeek-V2235.7BBF16263 GB
Solar-Open2-250B250.3BBF16280 GB
K-EXAONE-236B-A23B237.1BBF16265 GB
DeepSeek-V2-Chat235.7BBF16263 GB
bloom176.2BBF16197 GB
Qwen3-235B-A22B-Thinking-2507235.1BBF16263 GB
CYBER-FROST-3.8-BF16180.0BBF16201 GB

Vision-language models

6 models.

ModelParametersPublished asVRAM needed
Qwen3-VL-235B-A22B-Instruct235.7BBF16263 GB
Qwen3.8-Flash-Next180.0BBF16201 GB
Intern-S1240.7BBF16269 GB
Darwin-180B-RSI180.0BBF16201 GB
Qwen3.8-Flash-Next-Uncensored180.0BBF16201 GB
Huihui-Qwen3.8-Flash-Next-abliterated180.0BBF16201 GB

Fits at INT4

Models that fit only after quantizing to 4 bits (0.5 byte per parameter). Needs a quantized checkpoint actually published for the model; check its Hugging Face page before relying on a row.

Text models

16 models.

ModelParametersPublished asVRAM needed
Ornith-1.0-397B-FP8397.0BF8_E4M3222 GB
Llama-3.1-405B-FP8405.9BF8_E4M3227 GB
Qwen3-Coder-480B-A35B-Instruct-FP8480.2BF8_E4M3268 GB
Ornith-1.0-397B-FP8396.8BF8_E4M3222 GB
Ornith-1.5-397B403.4BBF16225 GB
Ornith-1.0-397B396.8BBF16222 GB
Ornith-1.5-397B-FP8403.4BF8_E4M3225 GB
Llama-3.1-405B405.9BBF16227 GB
GLM-4.5358.3BBF16200 GB
GLM-4.7358.3BBF16200 GB
Qwen3-Coder-480B-A35B-Instruct480.2BBF16268 GB
GLM-4.7-FP8358.5BF8_E4M3200 GB
MiniMax-Text-01-hf456.1BBF16255 GB
GLM-4.6356.8BBF16199 GB
Llama-3.1-405B-Instruct405.9BBF16227 GB
Trinity-Large-Preview398.6BBF16223 GB

Vision-language models

5 models.

ModelParametersPublished asVRAM needed
Qwen3.5-397B-A17B-FP8403.4BF8_E4M3225 GB
MiniMax-M3-MXFP8440.3BF8_E4M3246 GB
MiniMax-M3427.0BBF16239 GB
Qwen3.5-397B-A17B403.4BBF16225 GB
Llama-4-Maverick-17B-128E-Instruct-FP8401.6BF8_E4M3224 GB

Other models

1 model.

ModelParametersPublished asVRAM needed
Hermes-3-Llama-3.1-405B-FP8405.9BF8_E4M3227 GB

How these numbers are computed

Required VRAM is the weight size at each precision times a flat 1.2 overhead, the same figure every model page shows. It does not include a long context: the KV cache grows with every token, so a model near the top of a tier can need the next one at long context. Read how much VRAM you need for LLMs for the method, the VRAM and quantization glossary entries for the terms, and the VRAM calculator to size a model that is not listed.

Looking for a model by job rather than by memory? Start with coding, reasoning, chat and assistants, vision-language or see the full models directory.

Other VRAM tiers

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.