How to fine-tune Qwen3.8 (LoRA, QLoRA)

Checkpoints, thinking handling and license differences (Apache 2.0, qwen-community-1.0, qwen3.8-max) when fine-tuning Qwen3.8 models.

This guide covers adapting models from the Qwen3.8 family. Memory and GPU counts appear below. See also LoRA, QLoRA, the training cost calculator and fine-tuning.

Licenses differ by model

Check the license before you build on a checkpoint. The model cards list:

  • Qwen3.8-27B: Apache 2.0
  • Qwen3.8-Flash-Next: qwen-community-1.0 (a custom community license)
  • Qwen3.8-2.4T-A95B: qwen3.8-max (a custom license)

Review the full text of the two custom licenses for limits on derivatives and commercial use before training.

Which checkpoint to start from

The Qwen organization lists the three post-trained models and an FP8 variant of each. No Base checkpoints are listed. The FP8 build uses fine-grained FP8 quantization with block size 128, intended for serving. Train adapters on the unquantized repository (QLoRA quantizes on load).

Chat template and thinking

The models think by default. For Qwen3.8-27B and Flash-Next you can disable thinking with enable_thinking, and preserve_thinking keeps earlier reasoning in the history. Qwen3.8-2.4T-A95B always thinks, so training data for it should keep reasoning blocks. Use the shipped template in training and serving the same way, and decide whether your data carries reasoning traces.

Architecture notes

The 2.4T-A95B model is text-only and is a very large Mixture-of-Experts model, so full fine-tuning is rarely practical. The 27B model accepts images and video. The cards do not document LoRA target modules, so inspect module names in the checkpoint first.

Memory to fine-tune Qwen3.8, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Qwen3.8-2.4T-A95B2446.2B36451 GBNo live fit4556 GBNo live fit1175 GBNo live fit1.5 × 10^22 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.