How to run Kimi K3 with vLLM and SGLang

Serving Kimi K3: vLLM version, parser flags, trust-remote-code, reasoning_effort levels, vision input and preserved thinking in multi-turn use.

This guide covers serving the Kimi K3 family. Memory sizing and launch commands appear below.

Engine support

The model card recommends vLLM, SGLang and TokenSpeed, and points to a recipe or cookbook for each. The vLLM recipe states vLLM 0.29.0 or newer. The card gives no minimum version for SGLang.

Flags

The vLLM recipe uses:

  • --tool-call-parser kimi_k3
  • --reasoning-parser kimi_k3
  • --trust-remote-code

The card also says trust_remote_code=True is needed when loading with Transformers.

Reasoning

Reasoning is always on. Set effort with the reasoning_effort field: low, high or max (the default). The reply carries a reasoning_content field.

Multi-turn and tools

Kimi K3 uses preserved thinking. The card says to pass the full assistant message back as returned, including reasoning_content and tool_calls. Stripping them breaks continuity across tool use.

Context and vision

  • Context is 1,048,576 tokens. The card documents no rope or YaRN override.
  • The model is natively multimodal (text, images, video) with a MoonViT-V2 vision encoder. The vLLM recipe shows images sent as URLs through the OpenAI-compatible chat endpoint.
  • Weights are trained with MXFP4 weights and MXFP8 activations.

Sampling

The card lists temperature 1.0, with top-p 0.95 for single-step tasks and 1.0 for agentic tasks.

Pitfalls

  • Forgetting --trust-remote-code fails at load.
  • Dropping reasoning_content from history in agent loops.

Launch commands by Kimi K3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

No Kimi K3 model of a size we can compute is in the catalog yet. See the Kimi K3 model list for what is published.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.