How to fine-tune Qwen3.6 (LoRA, QLoRA)

Starting checkpoints, chat template and thinking handling, MoE and multimodal notes, and Apache 2.0 license terms for fine-tuning Qwen3.6.

This guide covers adapting models from the Qwen3.6 family. Memory and GPU counts appear below. See also LoRA, QLoRA, the training cost calculator and fine-tuning.

Which checkpoint to start from

The Qwen organization lists Qwen3.6-35B-A3B and Qwen3.6-27B, each with an FP8 variant. No Base checkpoint is listed for this generation. Start from the released weights, and train adapters on the unquantized repository rather than the FP8 one.

Chat template and thinking

The models think by default and emit <think> blocks. enable_thinking can be set false through chat_template_kwargs, and Qwen3.6-27B adds preserve_thinking to retain earlier reasoning in the history. Decide up front which format your training data uses and apply the shipped template the same way at train and inference time. If you train for agent use, matching the preserve_thinking setting you plan to serve with is sensible.

Architecture notes

Qwen3.6-35B-A3B is a Mixture-of-Experts model; Qwen3.6-27B is the other size. Both accept images and video. The cards do not document LoRA target modules, so inspect module names in the checkpoint. If you only want text behavior, keep vision components frozen. Use a recent Transformers release.

License

Both Qwen3.6-35B-A3B and Qwen3.6-27B are released under Apache 2.0, which permits commercial use and derivative works under its standard terms.

Memory to fine-tune Qwen3.6, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

No Qwen3.6 model of a size we can compute is in the catalog yet. See the Qwen3.6 model list for what is published.

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.