AI models that fit on a 288 GB GPU
Open models that fit in 288 GB of VRAM: 25 as published, 16 at FP8 and 22 at INT4. Each model is listed on the smallest tier it fits at that precision, so anything smaller is on the 192 GB page or below.
Fits as published
Models whose published weights, plus the flat overhead, fit this much memory with no quantization.
Text models
16 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| MiniMax-M2.7 | 228.7B | F8_E4M3 | 256 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 123.6B | BF16 | 276 GB |
| MiniMax-M2.5 | 228.7B | F8_E4M3 | 256 GB |
| MiniMax-M2 | 228.7B | F8_E4M3 | 256 GB |
| Qwen3-235B-A22B-Instruct-2507-FP8 | 235.1B | F8_E4M3 | 263 GB |
| GLM-4.5-Air | 110.5B | BF16 | 247 GB |
| Laguna-S-2.1 | 117.6B | BF16 | 263 GB |
| Qwen3-235B-A22B-Thinking-2507-FP8 | 235.1B | F8_E4M3 | 263 GB |
| MiniMax-M2.1 | 228.7B | F8_E4M3 | 256 GB |
| Qwen3-VL-235B-A22B-Instruct-FP8-dynamic | 235.8B | F8_E4M3 | 263 GB |
| DeepSeek-V2.5-1210-FP8 | 235.7B | F8_E4M3 | 263 GB |
| Ling-3.0-flash | 127.5B | BF16 | 285 GB |
| Qwen3-235B-A22B-FP8 | 235.1B | F8_E4M3 | 263 GB |
| Solar-Open-100B | 102.7B | BF16 | 229 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 | 123.6B | BF16 | 276 GB |
| GLM-4.5-Air-Base | 110.5B | BF16 | 247 GB |
Vision-language models
9 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3.5-122B-A10B | 125.1B | BF16 | 280 GB |
| Llama-4-Scout-17B-16E-Instruct | 108.6B | BF16 | 243 GB |
| Qwen3.8-Flash-Next-FP8 | 180.0B | F8_E4M3 | 201 GB |
| Qwen3-VL-235B-A22B-Instruct-FP8 | 235.7B | F8_E4M3 | 263 GB |
| Llama-3.2-90B-Vision-Instruct | 88.6B | BF16 | 198 GB |
| GLM-4.5V | 107.7B | BF16 | 241 GB |
| Qwen3.8-Flash-Next-FP8 | 180.0B | F8_E4M3 | 201 GB |
| Qwen3.8-Flash-Next-Uncensored-FP8 | 180.0B | F8_E4M3 | 201 GB |
| Qwen3.8-Flash-Next-UNCENSORED-FP8 | 180.0B | F8_E4M3 | 201 GB |
Fits at FP8
Models that fit only after quantizing the weights to 8 bits (1 byte per parameter). Needs an FP8 checkpoint or an engine that quantizes on load, and a GPU with FP8 support.
Text models
10 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3-235B-A22B | 235.1B | BF16 | 263 GB |
| Step-3.5-Flash | 199.4B | BF16 | 223 GB |
| Qwen3-235B-A22B-Instruct-2507 | 235.1B | BF16 | 263 GB |
| DeepSeek-V2 | 235.7B | BF16 | 263 GB |
| Solar-Open2-250B | 250.3B | BF16 | 280 GB |
| K-EXAONE-236B-A23B | 237.1B | BF16 | 265 GB |
| DeepSeek-V2-Chat | 235.7B | BF16 | 263 GB |
| bloom | 176.2B | BF16 | 197 GB |
| Qwen3-235B-A22B-Thinking-2507 | 235.1B | BF16 | 263 GB |
| CYBER-FROST-3.8-BF16 | 180.0B | BF16 | 201 GB |
Vision-language models
6 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3-VL-235B-A22B-Instruct | 235.7B | BF16 | 263 GB |
| Qwen3.8-Flash-Next | 180.0B | BF16 | 201 GB |
| Intern-S1 | 240.7B | BF16 | 269 GB |
| Darwin-180B-RSI | 180.0B | BF16 | 201 GB |
| Qwen3.8-Flash-Next-Uncensored | 180.0B | BF16 | 201 GB |
| Huihui-Qwen3.8-Flash-Next-abliterated | 180.0B | BF16 | 201 GB |
Fits at INT4
Models that fit only after quantizing to 4 bits (0.5 byte per parameter). Needs a quantized checkpoint actually published for the model; check its Hugging Face page before relying on a row.
Text models
16 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Ornith-1.0-397B-FP8 | 397.0B | F8_E4M3 | 222 GB |
| Llama-3.1-405B-FP8 | 405.9B | F8_E4M3 | 227 GB |
| Qwen3-Coder-480B-A35B-Instruct-FP8 | 480.2B | F8_E4M3 | 268 GB |
| Ornith-1.0-397B-FP8 | 396.8B | F8_E4M3 | 222 GB |
| Ornith-1.5-397B | 403.4B | BF16 | 225 GB |
| Ornith-1.0-397B | 396.8B | BF16 | 222 GB |
| Ornith-1.5-397B-FP8 | 403.4B | F8_E4M3 | 225 GB |
| Llama-3.1-405B | 405.9B | BF16 | 227 GB |
| GLM-4.5 | 358.3B | BF16 | 200 GB |
| GLM-4.7 | 358.3B | BF16 | 200 GB |
| Qwen3-Coder-480B-A35B-Instruct | 480.2B | BF16 | 268 GB |
| GLM-4.7-FP8 | 358.5B | F8_E4M3 | 200 GB |
| MiniMax-Text-01-hf | 456.1B | BF16 | 255 GB |
| GLM-4.6 | 356.8B | BF16 | 199 GB |
| Llama-3.1-405B-Instruct | 405.9B | BF16 | 227 GB |
| Trinity-Large-Preview | 398.6B | BF16 | 223 GB |
Vision-language models
5 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3.5-397B-A17B-FP8 | 403.4B | F8_E4M3 | 225 GB |
| MiniMax-M3-MXFP8 | 440.3B | F8_E4M3 | 246 GB |
| MiniMax-M3 | 427.0B | BF16 | 239 GB |
| Qwen3.5-397B-A17B | 403.4B | BF16 | 225 GB |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | 401.6B | F8_E4M3 | 224 GB |
Other models
1 model.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Hermes-3-Llama-3.1-405B-FP8 | 405.9B | F8_E4M3 | 227 GB |
How these numbers are computed
Required VRAM is the weight size at each precision times a flat 1.2 overhead, the same figure every model page shows. It does not include a long context: the KV cache grows with every token, so a model near the top of a tier can need the next one at long context. Read how much VRAM you need for LLMs for the method, the VRAM and quantization glossary entries for the terms, and the VRAM calculator to size a model that is not listed.
Looking for a model by job rather than by memory? Start with coding, reasoning, chat and assistants, vision-language or see the full models directory.