GLM-4.5 models

12 GLM-4.5 models on Hugging Face, from 10.3B to 358.5B parameters. At the precision each one is published in, the smallest needs about 23.0 GB of VRAM (GLM-4.6V-Flash, cheapest live fit: RTX A5000) and the largest about 801 GB (GLM-4.5). The cheapest way to run GLM-4.6V-Flash is $0.176/hr.

GLM-4.5 models

ModelParametersPublished asVRAM neededLive GPU fitEst. $/hr
GLM-4.6V-Flash10.3BBF1623.0 GBRTX A5000$0.176/hr
GLM-4.7-Flash31.2BBF1669.8 GBA100$1.21/hr
GLM-4.7-Flash31.2BBF1669.8 GBA100$1.21/hr
GLM-4.5V107.7BBF16241 GBRTX A6000 × 6$2.18/hr
GLM-4.5-Air110.5BBF16247 GBRTX A6000 × 6$2.18/hr
GLM-4.5-Air110.5BBF16247 GBRTX A6000 × 6$2.18/hr
GLM-4.5-Air-Base110.5BBF16247 GBRTX A6000 × 6$2.18/hr
GLM-4.5-Air-FP8110.5BF8_E4M3124 GBRTX 5060 Ti × 8$0.880/hr
GLM-4.6356.8BBF16797 GBNo live fit–
GLM-4.5358.3BBF16801 GBNo live fit–
GLM-4.7358.3BBF16801 GBNo live fit–
GLM-4.7-FP8358.5BF8_E4M3401 GBRTX PRO 6000 × 5$6.88/hr

VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

More from Z.ai

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.