How to run Kimi K3 with vLLM and SGLang
Serving Kimi K3: vLLM version, parser flags, trust-remote-code, reasoning_effort levels, vision input and preserved thinking in multi-turn use.
This guide covers serving the Kimi K3 family. Memory sizing and launch commands appear below.
Engine support
The model card recommends vLLM, SGLang and TokenSpeed, and points to a recipe or cookbook for each. The vLLM recipe states vLLM 0.29.0 or newer. The card gives no minimum version for SGLang.
Flags
The vLLM recipe uses:
--tool-call-parser kimi_k3--reasoning-parser kimi_k3--trust-remote-code
The card also says trust_remote_code=True is needed when loading with Transformers.
Reasoning
Reasoning is always on. Set effort with the reasoning_effort field: low, high or max (the default). The reply carries a reasoning_content field.
Multi-turn and tools
Kimi K3 uses preserved thinking. The card says to pass the full assistant message back as returned, including reasoning_content and tool_calls. Stripping them breaks continuity across tool use.
Context and vision
- Context is 1,048,576 tokens. The card documents no rope or YaRN override.
- The model is natively multimodal (text, images, video) with a MoonViT-V2 vision encoder. The vLLM recipe shows images sent as URLs through the OpenAI-compatible chat endpoint.
- Weights are trained with MXFP4 weights and MXFP8 activations.
Sampling
The card lists temperature 1.0, with top-p 0.95 for single-step tasks and 1.0 for agentic tasks.
Pitfalls
- Forgetting
--trust-remote-codefails at load. - Dropping
reasoning_contentfrom history in agent loops.
Launch commands by Kimi K3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
No Kimi K3 model of a size we can compute is in the catalog yet. See the Kimi K3 model list for what is published.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/moonshotai/Kimi-K3
- https://huggingface.co/moonshotai/Kimi-K3/raw/main/README.md
- https://recipes.vllm.ai/moonshotai/Kimi-K3
Updated 2026-10-07.