How to fine-tune Hy4 (LoRA and full)

Fine-tuning Tencent Hy4-preview: instruct-only checkpoints, LoRA defaults and target modules, the hy_v4 chat template and the router bias note.

Tencent ships a fine-tuning pipeline with the Hy4 family. For method and cost choices see fine-tuning and the training cost calculator.

Checkpoint

Only instruct models are listed: Hy4-preview and Hy4-preview-FP8. The card mentions no base checkpoint. Train from the unquantized instruct weights.

Tooling

The finetune guide offers three routes: DeepSpeed with the Hugging Face Trainer, LLaMA-Factory and ms-swift. Full fine-tuning and LoRA are both supported. The default LoRA setup is rank 64, alpha 128, dropout 0.05.

LoRA target modules

  • DeepSpeed route: q_proj, k_proj, v_proj, o_proj
  • LLaMA-Factory route: q_a_proj, q_b_proj, kv_a_proj_with_mqa, kv_b_proj, o_proj
  • ms-swift: set through its config

Chat template and data

Data is JSONL with message lists. The model uses the hy_v4 chat template, and reasoning_effort selects high (slow thinking) or no_think (fast). The guide says the system prompt for training and inference is empty by default, though you can customize it.

MoE note

The router's e_score_correction_bias is a buffer that the training script loads automatically. The guide says not to ignore a failure to load it.

License

The card lists Apache License 2.0, so derivatives carry no extra limits beyond that license.

Memory to fine-tune Hy4, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
Hy4-preview780.0B11622 GBNo live fit1453 GBNo live fit375 GBRTX A6000 × 8 · $2.90/hr4.7 × 10^21 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.