AI models that fit on a 141 GB GPU
Open models that fit in 141 GB of VRAM: 15 as published, 10 at FP8 and 30 at INT4. Each model is listed on the smallest tier it fits at that precision, so anything smaller is on the 96 GB page or below.
GPUs with 141 GB
Cards whose datasheet VRAM puts them in this tier, up to the next tier at 192 GB. Prices are the lowest live data-center on-demand rate per GPU.
Fits as published
Models whose published weights, plus the flat overhead, fit this much memory with no quantization.
Text models
13 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3-30B-A3B-abliterated | 30.5B | F32 | 136 GB |
| Kimi-Linear-48B-A3B-Instruct | 49.1B | BF16 | 110 GB |
| Llama-3_3-Nemotron-Super-49B-v1 | 49.9B | BF16 | 111 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-FP8 | 123.6B | F8_E4M3 | 138 GB |
| GLM-4.5-Air-FP8 | 110.5B | F8_E4M3 | 124 GB |
| Laguna-S-2.1-FP8 | 117.6B | F8_E4M3 | 131 GB |
| Nous-Hermes-2-Mixtral-8x7B-DPO | 46.7B | BF16 | 104 GB |
| Llama-3_3-Nemotron-Super-49B-v1_5 | 49.9B | BF16 | 111 GB |
| Nemotron-H-56B-Base-8K | 56.3B | BF16 | 126 GB |
| Qwen2-57B-A14B-Instruct | 57.4B | BF16 | 128 GB |
| Jamba-v0.1 | 51.6B | BF16 | 115 GB |
| gemma-2-27b | 27.2B | F32 | 122 GB |
| Qwen2-57B-A14B | 57.4B | BF16 | 128 GB |
Vision-language models
1 model.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3.5-122B-A10B-FP8 | 125.1B | F8_E4M3 | 140 GB |
Fits at FP8
Models that fit only after quantizing the weights to 8 bits (1 byte per parameter). Needs an FP8 checkpoint or an engine that quantizes on load, and a GPU with FP8 support.
Text models
6 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 123.6B | BF16 | 138 GB |
| GLM-4.5-Air | 110.5B | BF16 | 123 GB |
| Laguna-S-2.1 | 117.6B | BF16 | 131 GB |
| Solar-Open-100B | 102.7B | BF16 | 115 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 | 123.6B | BF16 | 138 GB |
| GLM-4.5-Air-Base | 110.5B | BF16 | 123 GB |
Vision-language models
4 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3.5-122B-A10B | 125.1B | BF16 | 140 GB |
| Llama-4-Scout-17B-16E-Instruct | 108.6B | BF16 | 121 GB |
| Llama-3.2-90B-Vision-Instruct | 88.6B | BF16 | 99.0 GB |
| GLM-4.5V | 107.7B | BF16 | 120 GB |
Fits at INT4
Models that fit only after quantizing to 4 bits (0.5 byte per parameter). Needs a quantized checkpoint actually published for the model; check its Hugging Face page before relying on a row.
Text models
19 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| MiniMax-M2.7 | 228.7B | F8_E4M3 | 128 GB |
| MiniMax-M2.5 | 228.7B | F8_E4M3 | 128 GB |
| Qwen3-235B-A22B | 235.1B | BF16 | 131 GB |
| MiniMax-M2 | 228.7B | F8_E4M3 | 128 GB |
| Qwen3-235B-A22B-Instruct-2507-FP8 | 235.1B | F8_E4M3 | 131 GB |
| Step-3.5-Flash | 199.4B | BF16 | 111 GB |
| Qwen3-235B-A22B-Instruct-2507 | 235.1B | BF16 | 131 GB |
| Qwen3-235B-A22B-Thinking-2507-FP8 | 235.1B | F8_E4M3 | 131 GB |
| DeepSeek-V2 | 235.7B | BF16 | 132 GB |
| MiniMax-M2.1 | 228.7B | F8_E4M3 | 128 GB |
| Qwen3-VL-235B-A22B-Instruct-FP8-dynamic | 235.8B | F8_E4M3 | 132 GB |
| Solar-Open2-250B | 250.3B | BF16 | 140 GB |
| DeepSeek-V2.5-1210-FP8 | 235.7B | F8_E4M3 | 132 GB |
| K-EXAONE-236B-A23B | 237.1B | BF16 | 132 GB |
| DeepSeek-V2-Chat | 235.7B | BF16 | 132 GB |
| bloom | 176.2B | BF16 | 98.5 GB |
| Qwen3-235B-A22B-FP8 | 235.1B | F8_E4M3 | 131 GB |
| Qwen3-235B-A22B-Thinking-2507 | 235.1B | BF16 | 131 GB |
| CYBER-FROST-3.8-BF16 | 180.0B | BF16 | 101 GB |
Vision-language models
11 models.
| Model | Parameters | Published as | VRAM needed |
|---|---|---|---|
| Qwen3-VL-235B-A22B-Instruct | 235.7B | BF16 | 132 GB |
| Qwen3.8-Flash-Next | 180.0B | BF16 | 101 GB |
| Qwen3.8-Flash-Next-FP8 | 180.0B | F8_E4M3 | 101 GB |
| Qwen3-VL-235B-A22B-Instruct-FP8 | 235.7B | F8_E4M3 | 132 GB |
| Intern-S1 | 240.7B | BF16 | 135 GB |
| Qwen3.8-Flash-Next-FP8 | 180.0B | F8_E4M3 | 101 GB |
| Qwen3.8-Flash-Next-Uncensored-FP8 | 180.0B | F8_E4M3 | 101 GB |
| Darwin-180B-RSI | 180.0B | BF16 | 101 GB |
| Qwen3.8-Flash-Next-Uncensored | 180.0B | BF16 | 101 GB |
| Huihui-Qwen3.8-Flash-Next-abliterated | 180.0B | BF16 | 101 GB |
| Qwen3.8-Flash-Next-UNCENSORED-FP8 | 180.0B | F8_E4M3 | 101 GB |
How these numbers are computed
Required VRAM is the weight size at each precision times a flat 1.2 overhead, the same figure every model page shows. It does not include a long context: the KV cache grows with every token, so a model near the top of a tier can need the next one at long context. Read how much VRAM you need for LLMs for the method, the VRAM and quantization glossary entries for the terms, and the VRAM calculator to size a model that is not listed.
Looking for a model by job rather than by memory? Start with coding, reasoning, chat and assistants, vision-language or see the full models directory.