Kimi K2.5
A 1T-parameter (32B active) native multimodal agentic model. View on Hugging Face.
What Kimi K2.5 is
Kimi K2.5 is Moonshot AI's native multimodal agentic model, a 1T-parameter mixture-of-experts model (32B active) continually pretrained on approximately 15 trillion mixed vision-language tokens atop Kimi-K2-Base. It integrates vision and language understanding with agentic tool use, instant and thinking modes, and an "Agent Swarm" mode that decomposes tasks across dynamically instantiated sub-agents. Its weights publish in Moonshot's native INT4 quantization, the same scheme as Kimi-K2-Thinking.
License note: MIT terms, with one modification: a commercial product or service with more than 100 million monthly active users or more than $20 million/month revenue must prominently display "Kimi K2.5" in its user interface. Facts on this page are sourced from Kimi K2.5's Hugging Face model card, not benchmarked by Aquanode.
What it's used for
- Vision-language agentic tasks
- Coding from visual specifications (UI designs)
- Multi-agent tool-use workflows
Benchmarks (published by Moonshot AI)
Published by Moonshot AI, not measured by Aquanode.
Deploy Kimi K2.5
Needs 6× AMD MI300X at its published (Native INT4) precision. Aquanode doesn't run a hosted inference API for Kimi K2.5. Spin up a GPU pod sized for it below and install your own inference stack (vLLM, SGLang, TGI, llama.cpp, Ollama, or whatever you prefer) on it.
Cheapest current fit at Kimi K2.5's published (Native INT4) precision: 6× AMD MI300X, at $2.39/hr per GPU ($14.34/hr total).
More models
DeepSeek-R1-0528
Reasoning
Llama 3.3 70B Instruct
LLM
Qwen2.5 7B Instruct
LLM
FLUX.1 [dev]
Image generation
deepseek-coder-6.7b-instruct
Coding
DeepSeek-OCR
Vision
DeepSeek-R1-Distill-Llama-70B
Reasoning
DeepSeek-R1-Distill-Llama-8B-abliterated
Reasoning
DeepSeek-R1-Distill-Qwen-1.5B
Reasoning
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Reasoning
DeepSeek-V3
LLM
DeepSeek-V3-0324
LLM
DeepSeek-V3.1
LLM
DeepSeek-V3.2-Exp
LLM
Dia-1.6B
Speech / TTS
Dolphin3.0-R1-Mistral-24B
Reasoning
EXAONE-3.5-32B-Instruct
LLM
FLUX.1-Kontext-dev
Image generation
FLUX.1-schnell
Image generation
FLUX.2-dev
Image generation
gemma-3-12b-it
Vision
gemma-3-4b-it
Vision
gemma-4-31B-it
Vision
GLM-4.5-Air
LLM
GLM-4.5V
Vision
GLM-4.6
LLM
GLM-4.7-Flash
LLM
GLM-5
LLM
GLM-5.1
LLM
GLM-5.2
LLM
granite-3.2-8b-instruct
LLM
Hermes-3-Llama-3.1-405B-FP8
LLM
Hermes-3-Llama-3.1-70B
LLM
Hermes-3-Llama-3.1-8B
LLM
HiDream-I1-Full
Image generation
InternVL3-78B
Vision
Kimi-K2-Instruct
LLM
Kimi-K2-Instruct-0905
LLM
Krea-2-Raw
Image generation
turn-detector
Speech / TTS
Llama-3.1-405B-Instruct
LLM
Meta-Llama-3.1-70B-Instruct-FP8
LLM
Llama-3.1-8B-Instruct
LLM
Llama-3.1-8B-Lexi-Uncensored-V2
LLM
Llama-3.1-Nemotron-Nano-8B-v1
LLM
Llama-3.2-3B-Instruct
LLM
Llama-4-Maverick-17B-128E-Instruct-FP8
Vision
Llama-4-Scout-17B-16E-Instruct
Vision
Llama-Guard-3-8B
LLM
Llama-Guard-4-12B
LLM
LTX-2
Video generation
MiniMax-H3
LLM
MiniMax-M2.5
LLM
MiniMax-M2.7
LLM
Mistral-Small-24B-Instruct-2501
LLM
Mixtral-8x7B-Instruct-v0.1
LLM
mochi-1-preview
Video generation
multilingual-e5-large-instruct
Embedding
Nanonets-OCR-s
Vision
Nous-Hermes-2-Mistral-7B-DPO
LLM
orpheus-3b-0.1-ft
Speech / TTS
parakeet-tdt-0.6b-v3
Speech / TTS
Phi-3.5-mini-instruct
LLM
Phi-3-mini-4k-instruct
LLM
phi-4
LLM
Qwen-Image
Image generation
Qwen2.5-14B-Instruct
LLM
Qwen2.5-3B-Instruct
LLM
Qwen2.5-Coder-32B-Instruct
Coding
Qwen3-1.7B
LLM
Qwen3-1.7B-Base
LLM
Qwen3-14B-Base
LLM
Qwen3-235B-A22B-Instruct-2507-FP8
LLM
Qwen3-235B-A22B-Thinking-2507
Reasoning
Qwen3-30B-A3B
LLM
Qwen3-32B
LLM
Qwen3-4B
LLM
Qwen3-4B-Base
LLM
Qwen3.5-122B-A10B
LLM
Qwen3.5-397B-A17B
LLM
Qwen3.5-9B
LLM
Qwen3.6-35B-A3B
LLM
Qwen3.8-27B-FP8
LLM
Qwen3-8B
LLM
Qwen3-Coder-480B-A35B-Instruct
Coding
Qwen3-Coder-Next
Coding
Qwen3-Embedding-0.6B
Embedding
Qwen3-Next-80B-A3B-Instruct
Reasoning
Qwen3-Next-80B-A3B-Thinking
Reasoning
Qwen3-VL-32B-Instruct
Vision
ReaderLM-v2
LLM
RealVisXL_V5.0
Image generation
SmolLM2-360M-Instruct
LLM
stable-diffusion-xl-base-1.0
Image generation
Trinity-Large-Preview
LLM
Wayfarer-12B
LLM
whisper-large-v3
Speech / TTS
gpt-oss-120b
Reasoning
gpt-oss-20b
Reasoning
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
LLM
MiniMax-M2
Coding
DeepSeek-V4-Pro-0813
Reasoning
DeepSeek-V4-Flash-0731
Reasoning
Wan2.2-T2V-A14B
Video generation
Wan2.2-I2V-A14B
Video generation
Wan2.2-TI2V-5B
Video generation
gemma-4-12B-it
Vision
gemma-4-26B-A4B-it
Vision
gemma-4-E4B-it
Vision
gemma-4-E2B-it
Vision