Fine-tuning Kimi K3: checkpoints and license limits

What the Kimi K3 card and license say about fine-tuning: checkpoint availability, trust-remote-code, preserved thinking format and revenue limits.

Documentation for adapting the Kimi K3 family is thin. The model card gives no fine-tuning guide, so this page only records what is stated. For general method choices see fine-tuning and the training cost calculator.

Checkpoint

The card lists a single release and mentions no base checkpoint. The weights were trained with quantization-aware training (MXFP4 weights, MXFP8 activations) from the SFT stage onward, so check what precision your trainer loads. If you use LoRA or QLoRA, confirm the trainer supports this quantized format.

Loading

trust_remote_code=True is required for both pipeline and direct model loading. The model has a MoonViT-V2 vision encoder, so decide up front whether to train text only or with image inputs.

Chat format

Reasoning is always enabled, with reasoning_content returned and reasoning_effort of low, high or max. The card requires full assistant messages, including reasoning_content and tool_calls, to be passed back in multi-turn conversations. Build training examples the same way so the model sees the format it will be served with.

License

Weights and code use the Kimi K3 License, a custom license rather than MIT or Apache. Fine-tuning and derivative works are permitted. Read these sections in the LICENSE file:

  • Section 2: if your aggregate revenue exceeds $20 million over any consecutive 12 months, a separate commercial agreement with Moonshot AI is required for commercial use.
  • Section 3: very large products must display "Kimi K3" on the interface.
  • Section 4: internal use and use through official channels are exempt from those commercial restrictions.

Confirm the exact wording against the LICENSE file before shipping a derivative.

Memory to fine-tune Kimi K3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

No Kimi K3 model of a size we can compute is in the catalog yet. See the Kimi K3 model list for what is published.

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.