Kimi models, generation by generation

Moonshot AI's Kimi open-weight line: K2 (1T MoE), K2.5, K3 (2.8T) and Kimi-VL. Sizes, context, licenses, and what to use today.

What the Kimi line is

Kimi is the model family from Moonshot AI, published under the moonshotai organization on Hugging Face. Its two main generations are both very large mixture-of-experts (MoE) models built for agentic coding and long-horizon tool use. K2 is a 1 trillion parameter model, and K3 is 2.8 trillion. A separate small vision-language line, Kimi-VL, is the exception.

Generations in order

Kimi K2 (July 2025, per the Hugging Face repository date). K2 is a MoE language model with 1 trillion total and 32 billion active parameters. Its card lists 384 experts with 8 selected per token plus 1 shared expert, a 128K token context window, and training with the Muon optimizer. The license is a modified MIT license.

The hub holds several later updates:

  • K2-Instruct-0905 (September 2025) raised the context window from 128K to 256K tokens, with the same 1 trillion and 32 billion parameter sizes.
  • K2.5 (repository created January 2026) is described as a native multimodal agentic model, built by continual pretraining on about 15 trillion mixed visual and text tokens on top of Kimi-K2-Base. It keeps the 1T/32B MoE layout and the 256K context window, and the modified MIT license.

The hub also holds Kimi-K2-Base, the pretrained base that K2.5 builds on.

Kimi K3 (repository created June 2026). K3 is a new architecture, not a scaled K2. Its card gives 2.8 trillion total and 104 billion active parameters, 93 layers of which 69 use Kimi Delta Attention (a linear-attention variant) and 24 use gated multi-head latent attention, and a MoE that activates 16 of 896 experts. It is natively multimodal for text, images and video, with a 1 million token context window. Moonshot states about a 2.5x improvement in scaling efficiency over K2, which is a vendor-reported figure.

The K3 license is its own, not modified MIT. In short, it grants broad rights to use, modify, fine-tune and redistribute, with two conditions: a business that offers "Model as a Service" and earns more than 20 million US dollars over any 12 months must sign a separate agreement with Moonshot for commercial use, and products above 100 million monthly active users or 20 million US dollars in monthly revenue must display "Kimi K3" in the interface. Internal use is exempt. Read the full text before relying on this summary.

Both K2 and K3 are served with the same pattern in mind: very few experts fire per token (8 of 384 for K2, 16 of 896 for K3), so compute per token stays far below what the total parameter counts suggest. Memory does not shrink the same way, because every expert must be resident. When comparing K2.5 and K3, treat the jump as a change of architecture, context length and license, not only of size.

Side line: Kimi-VL

Kimi-VL (April 2025) is a small vision-language MoE with 16 billion total and about 3 billion active parameters (the card says 2.8B), a 128K token context window, and an MIT license. The card recommends Kimi-VL-A3B-Instruct for general perception, OCR, long video and long documents, and the separate Thinking variant for harder math and reasoning.

Which generation to use today

  • One GPU, images or documents: Kimi-VL-A3B-Instruct. It is the only model here with a small total size, and it is MIT licensed.
  • Large open text model with a familiar license: Kimi K2-Instruct-0905 or K2.5, both modified MIT, with 256K context. Pick K2.5 if you need image input; pick 0905 if text is enough.
  • Longest context and video input: K3, at 1 million tokens, if you have the multi-node capacity for 2.8 trillion parameters and the license terms above suit you.
  • Original K2 (128K): only if you need to match an existing deployment, since 0905 has the larger window at the same size.

Any of these can run on Aquanode GPUs.

Kimi generations

Every Kimi generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
Kimi K253.0B to 1026.9B1128 GBNo live fit–
Kimi K332.2B to 3.6B5.0 GBRTX 4070 Super$0.121/hr

Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.

Other Kimi lines

Side lines from Moonshot AI: specialised or companion models outside the main Kimi generations.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
Kimi VL116.4B36.7 GBRTX A6000$0.363/hr

Run and fine-tune Kimi

Per-generation guides: engine support, chat template and context flags, then fine-tune memory for LoRA and QLoRA.

Kimi compared with other models

Side-by-side VRAM, context length and license for models of a similar size.

Best models by task

Kimi models appear on these ranked task pages.

Sources

Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.