GLM-4.5 models
12 GLM-4.5 models on Hugging Face, from 10.3B to 358.5B parameters. At the precision each one is published in, the smallest needs about 23.0 GB of VRAM (GLM-4.6V-Flash, cheapest live fit: RTX A5000) and the largest about 801 GB (GLM-4.5). The cheapest way to run GLM-4.6V-Flash is $0.176/hr.
GLM-4.5 models
| Model | Parameters | Published as | VRAM needed | Live GPU fit | Est. $/hr |
|---|---|---|---|---|---|
| GLM-4.6V-Flash | 10.3B | BF16 | 23.0 GB | RTX A5000 | $0.176/hr |
| GLM-4.7-Flash | 31.2B | BF16 | 69.8 GB | A100 | $1.21/hr |
| GLM-4.7-Flash | 31.2B | BF16 | 69.8 GB | A100 | $1.21/hr |
| GLM-4.5V | 107.7B | BF16 | 241 GB | RTX A6000 × 6 | $2.18/hr |
| GLM-4.5-Air | 110.5B | BF16 | 247 GB | RTX A6000 × 6 | $2.18/hr |
| GLM-4.5-Air | 110.5B | BF16 | 247 GB | RTX A6000 × 6 | $2.18/hr |
| GLM-4.5-Air-Base | 110.5B | BF16 | 247 GB | RTX A6000 × 6 | $2.18/hr |
| GLM-4.5-Air-FP8 | 110.5B | F8_E4M3 | 124 GB | RTX 5060 Ti × 8 | $0.880/hr |
| GLM-4.6 | 356.8B | BF16 | 797 GB | No live fit | – |
| GLM-4.5 | 358.3B | BF16 | 801 GB | No live fit | – |
| GLM-4.7 | 358.3B | BF16 | 801 GB | No live fit | – |
| GLM-4.7-FP8 | 358.5B | F8_E4M3 | 401 GB | RTX PRO 6000 × 5 | $6.88/hr |
VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.