AI models that fit on a 192 GB GPU

Open models that fit in 192 GB of VRAM: 37 as published, 2 at FP8 and 17 at INT4. Each model is listed on the smallest tier it fits at that precision, so anything smaller is on the 141 GB page or below.

GPUs with 192 GB

Cards whose datasheet VRAM puts them in this tier, up to the next tier at 288 GB. Prices are the lowest live data-center on-demand rate per GPU.

GPUVRAMFrom per GPU hour
AMD MI300X192 GB$2.63/hr

Fits as published

Models whose published weights, plus the flat overhead, fit this much memory with no quantization.

Text models

31 models.

ModelParametersPublished asVRAM needed
Qwen-72B72.3BBF16162 GB
Llama-3.3-70B-Instruct70.6BBF16158 GB
Llama-3.1-70B-Instruct70.6BBF16158 GB
Qwen3-Coder-Next79.7BBF16178 GB
Qwen2.5-72B-Instruct72.7BBF16163 GB
Qwen3-Next-80B-A3B-Instruct81.3BBF16182 GB
Meta-Llama-3-70B70.6BBF16158 GB
DeepSeek-R1-Distill-Llama-70B70.6BBF16158 GB
Llama-3.1-70B-LatamGPT-SFT-1.070.6BBF16158 GB
Meta-Llama-3-70B-Instruct70.6BBF16158 GB
Hunyuan-A13B-Instruct80.4BBF16180 GB
Qwen2.5-72B72.7BBF16163 GB
Llama-3.1-70B70.6BBF16158 GB
Qwen3-Next-80B-A3B-Thinking81.3BBF16182 GB
Qwen2-72B-Instruct72.7BBF16163 GB
EXAONE-3.5-32B-Instruct32.0BF32143 GB
Meta-Llama-3.1-70B70.6BBF16158 GB
Meta-Llama-3.1-70B-Instruct70.6BBF16158 GB
Apertus-70B-Instruct-250970.6BBF16158 GB
Le_Triomphant-ECE-TW372.3BBF16162 GB
TW3-JRGL-v272.3BBF16162 GB
LongCat-Flash-Lite69.1BBF16154 GB
sarvam-30b32.2BF32144 GB
Meta-Llama-3.1-70B-Instruct70.6BBF16158 GB
Llama-3.1-Nemotron-70B-Instruct-HF70.6BBF16158 GB
Qwen2-72B72.7BBF16163 GB
Qwen1.5-72B-Chat72.3BBF16162 GB
Qwen1.5-72B72.3BBF16162 GB
AliceAI-Foundation-80B-A3B-Base81.3BBF16182 GB
Hermes-3-Llama-3.1-70B70.6BBF16158 GB
Kolibri-1-BF1678.1BBF16175 GB

Vision-language models

3 models.

ModelParametersPublished asVRAM needed
Qwen2.5-VL-72B-Instruct73.4BBF16164 GB
InternVL3-78B78.4BBF16175 GB
Qwen3.8-Flash-Next-Uncensored-MLX71.3BBF16159 GB

Image generation models

1 model.

ModelParametersPublished asVRAM needed
Cosmos3-Super-Text2Image-4Step64.0BBF16143 GB

Video generation models

1 model.

ModelParametersPublished asVRAM needed
Cosmos3-Super-Image2Video64.6BBF16144 GB

Other models

1 model.

ModelParametersPublished asVRAM needed
Isaac-0.535.7BF32160 GB

Fits at FP8

Models that fit only after quantizing the weights to 8 bits (1 byte per parameter). Needs an FP8 checkpoint or an engine that quantizes on load, and a GPU with FP8 support.

Text models

2 models.

ModelParametersPublished asVRAM needed
Ling-3.0-flash127.5BBF16142 GB
xLAM-8x22b-r140.6BBF16157 GB

Fits at INT4

Models that fit only after quantizing to 4 bits (0.5 byte per parameter). Needs a quantized checkpoint actually published for the model; check its Hugging Face page before relying on a row.

Text models

9 models.

ModelParametersPublished asVRAM needed
MiMo-V2.5310.8BF8_E4M3174 GB
MiMo-V2-Flash309.8BF8_E4M3173 GB
Hy3-preview298.8BBF16167 GB
Hy3-FP8298.8BF8_E4M3167 GB
Hy3298.8BBF16167 GB
GLM-5.3-Flash-FP8321.3BF8_E4M3180 GB
GLM-5.3-Flash-Uncensored-FP8321.3BF8_E4M3180 GB
IQuest-Q1320.3BBF16179 GB
GLM-5.3-Flash321.3BBF16180 GB

Vision-language models

6 models.

ModelParametersPublished asVRAM needed
GLM-5.3-Flash321.3BF8_E4M3180 GB
Inkling-Small266.0BBF16149 GB
step3321.0BBF16179 GB
GLM-5.3-Flash-BF16321.3BBF16180 GB
apex-flash-1321.3BBF16180 GB
apex-flash-1-abliterated321.3BBF16180 GB

Other models

2 models.

ModelParametersPublished asVRAM needed
GLM-5.3-Flash-UNCENSORED-FP8321.3BF8_E4M3180 GB
GLM-5.3-Flash-ABLITERATED-FP8321.3BF8_E4M3180 GB

How these numbers are computed

Required VRAM is the weight size at each precision times a flat 1.2 overhead, the same figure every model page shows. It does not include a long context: the KV cache grows with every token, so a model near the top of a tier can need the next one at long context. Read how much VRAM you need for LLMs for the method, the VRAM and quantization glossary entries for the terms, and the VRAM calculator to size a model that is not listed.

Looking for a model by job rather than by memory? Start with coding, reasoning, chat and assistants, vision-language or see the full models directory.

Other VRAM tiers

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.