How to fine-tune Qwen3.5 (LoRA, QLoRA)

Base versus post-trained Qwen3.5 checkpoints, chat template and thinking handling, multimodal notes and Apache 2.0 license terms.

This guide covers adapting models from the Qwen3.5 family. VRAM and GPU counts appear below. See also LoRA, QLoRA, the training cost calculator and fine-tuning.

Which checkpoint to start from

  • The main repositories (for example Qwen3.5-35B-A3B) are post-trained.
  • Base checkpoints are published for several sizes: the Qwen organization lists Qwen3.5-0.8B-Base, 2B-Base, 4B-Base, 9B-Base and 35B-A3B-Base.
  • Start from Base for domain adaptation or your own instruction tuning. Start from the post-trained model to keep its chat and reasoning behavior.

Chat template and thinking

The post-trained models think by default (<think> blocks) and enable_thinking can be set to false through chat_template_kwargs. Train with the shipped template and keep the thinking format consistent across your dataset. If your data has no reasoning traces, train in the non-thinking format.

Architecture notes

The 35B-A3B model is a Mixture-of-Experts model, and the family is a unified vision-language design trained with early fusion. The card does not document LoRA target modules or adapter settings, so inspect the checkpoint's module names. If you only want text behavior, keep the vision components frozen.

For libraries, the card points to Transformers main at release time, so use a recent version.

License

Qwen3.5-35B-A3B is released under Apache 2.0, which permits commercial use and derivative works under its standard terms. Confirm the license field on any other size you use.

Memory to fine-tune Qwen3.5, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Qwen-AgentWorld-35B-A3B34.7B516 GBRTX PRO 6000 × 6 · $8.25/hr64.6 GBA100 · $1.21/hr16.7 GBRTX A5000 · $0.176/hr2.1 × 10^20 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.