Deploy AI models on H100, A100, H200, and AMD MI300X GPUs with up to 40% cost savings. Lightning-fast machine learning inference on enterprise GPU infrastructure.
46 Qwen3.5 models on Hugging Face, from 556M to 403.4B parameters. At the precision each one is published in, the smallest needs about 1.2 GB of VRAM (gepard-1.0, cheapest live fit: RTX 5060 Ti) and the largest about 902 GB (Qwen3.5-397B-A17B). The cheapest way to run gepard-1.0 is $0.110/hr.
VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.