Kimi K2 models

5 Kimi K2 models on Hugging Face, from 3.0B to 1026.9B parameters. At the precision each one is published in, the smallest needs about 6.7 GB of VRAM (kimi-k2.6-eagle3-mla, cheapest live fit: RTX 5060 Ti) and the largest about 1147 GB (Kimi-K2-Instruct-0905). The cheapest way to run kimi-k2.6-eagle3-mla is $0.110/hr.

Kimi K2 models

ModelParametersPublished asVRAM neededLive GPU fitEst. $/hr
kimi-k2.6-eagle3-mla3.0BBF166.7 GBRTX 5060 Ti$0.110/hr
Kimi-K2-Instruct1026.4BF8_E4M31147 GBNo live fit–
Kimi-K2-Base1026.5BF8_E4M31147 GBNo live fit–
Kimi-K2-Instruct-09051026.5BF8_E4M31147 GBNo live fit–
Kimi-K2.51026.9BNative INT41128 GBNo live fit–

VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

More from Moonshot AI

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.