AI models that fit on a 141 GB GPU

Open models that fit in 141 GB of VRAM: 15 as published, 10 at FP8 and 30 at INT4. Each model is listed on the smallest tier it fits at that precision, so anything smaller is on the 96 GB page or below.

GPUs with 141 GB

Cards whose datasheet VRAM puts them in this tier, up to the next tier at 192 GB. Prices are the lowest live data-center on-demand rate per GPU.

GPUVRAMFrom per GPU hour
H200141 GB$3.98/hr
B200180 GB$7.47/hr

Fits as published

Models whose published weights, plus the flat overhead, fit this much memory with no quantization.

Text models

13 models.

ModelParametersPublished asVRAM needed
Qwen3-30B-A3B-abliterated30.5BF32136 GB
Kimi-Linear-48B-A3B-Instruct49.1BBF16110 GB
Llama-3_3-Nemotron-Super-49B-v149.9BBF16111 GB
NVIDIA-Nemotron-3-Super-120B-A12B-FP8123.6BF8_E4M3138 GB
GLM-4.5-Air-FP8110.5BF8_E4M3124 GB
Laguna-S-2.1-FP8117.6BF8_E4M3131 GB
Nous-Hermes-2-Mixtral-8x7B-DPO46.7BBF16104 GB
Llama-3_3-Nemotron-Super-49B-v1_549.9BBF16111 GB
Nemotron-H-56B-Base-8K56.3BBF16126 GB
Qwen2-57B-A14B-Instruct57.4BBF16128 GB
Jamba-v0.151.6BBF16115 GB
gemma-2-27b27.2BF32122 GB
Qwen2-57B-A14B57.4BBF16128 GB

Vision-language models

1 model.

ModelParametersPublished asVRAM needed
Qwen3.5-122B-A10B-FP8125.1BF8_E4M3140 GB

Other models

1 model.

ModelParametersPublished asVRAM needed
Mixtral-8x7B-Instruct-v0.146.7BBF16104 GB

Fits at FP8

Models that fit only after quantizing the weights to 8 bits (1 byte per parameter). Needs an FP8 checkpoint or an engine that quantizes on load, and a GPU with FP8 support.

Text models

6 models.

ModelParametersPublished asVRAM needed
NVIDIA-Nemotron-3-Super-120B-A12B-BF16123.6BBF16138 GB
GLM-4.5-Air110.5BBF16123 GB
Laguna-S-2.1117.6BBF16131 GB
Solar-Open-100B102.7BBF16115 GB
NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16123.6BBF16138 GB
GLM-4.5-Air-Base110.5BBF16123 GB

Vision-language models

4 models.

ModelParametersPublished asVRAM needed
Qwen3.5-122B-A10B125.1BBF16140 GB
Llama-4-Scout-17B-16E-Instruct108.6BBF16121 GB
Llama-3.2-90B-Vision-Instruct88.6BBF1699.0 GB
GLM-4.5V107.7BBF16120 GB

Fits at INT4

Models that fit only after quantizing to 4 bits (0.5 byte per parameter). Needs a quantized checkpoint actually published for the model; check its Hugging Face page before relying on a row.

Text models

19 models.

ModelParametersPublished asVRAM needed
MiniMax-M2.7228.7BF8_E4M3128 GB
MiniMax-M2.5228.7BF8_E4M3128 GB
Qwen3-235B-A22B235.1BBF16131 GB
MiniMax-M2228.7BF8_E4M3128 GB
Qwen3-235B-A22B-Instruct-2507-FP8235.1BF8_E4M3131 GB
Step-3.5-Flash199.4BBF16111 GB
Qwen3-235B-A22B-Instruct-2507235.1BBF16131 GB
Qwen3-235B-A22B-Thinking-2507-FP8235.1BF8_E4M3131 GB
DeepSeek-V2235.7BBF16132 GB
MiniMax-M2.1228.7BF8_E4M3128 GB
Qwen3-VL-235B-A22B-Instruct-FP8-dynamic235.8BF8_E4M3132 GB
Solar-Open2-250B250.3BBF16140 GB
DeepSeek-V2.5-1210-FP8235.7BF8_E4M3132 GB
K-EXAONE-236B-A23B237.1BBF16132 GB
DeepSeek-V2-Chat235.7BBF16132 GB
bloom176.2BBF1698.5 GB
Qwen3-235B-A22B-FP8235.1BF8_E4M3131 GB
Qwen3-235B-A22B-Thinking-2507235.1BBF16131 GB
CYBER-FROST-3.8-BF16180.0BBF16101 GB

Vision-language models

11 models.

ModelParametersPublished asVRAM needed
Qwen3-VL-235B-A22B-Instruct235.7BBF16132 GB
Qwen3.8-Flash-Next180.0BBF16101 GB
Qwen3.8-Flash-Next-FP8180.0BF8_E4M3101 GB
Qwen3-VL-235B-A22B-Instruct-FP8235.7BF8_E4M3132 GB
Intern-S1240.7BBF16135 GB
Qwen3.8-Flash-Next-FP8180.0BF8_E4M3101 GB
Qwen3.8-Flash-Next-Uncensored-FP8180.0BF8_E4M3101 GB
Darwin-180B-RSI180.0BBF16101 GB
Qwen3.8-Flash-Next-Uncensored180.0BBF16101 GB
Huihui-Qwen3.8-Flash-Next-abliterated180.0BBF16101 GB
Qwen3.8-Flash-Next-UNCENSORED-FP8180.0BF8_E4M3101 GB

How these numbers are computed

Required VRAM is the weight size at each precision times a flat 1.2 overhead, the same figure every model page shows. It does not include a long context: the KV cache grows with every token, so a model near the top of a tier can need the next one at long context. Read how much VRAM you need for LLMs for the method, the VRAM and quantization glossary entries for the terms, and the VRAM calculator to size a model that is not listed.

Looking for a model by job rather than by memory? Start with coding, reasoning, chat and assistants, vision-language or see the full models directory.

Other VRAM tiers

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.