How to fine-tune Hy4 (LoRA and full)
Fine-tuning Tencent Hy4-preview: instruct-only checkpoints, LoRA defaults and target modules, the hy_v4 chat template and the router bias note.
Tencent ships a fine-tuning pipeline with the Hy4 family. For method and cost choices see fine-tuning and the training cost calculator.
Checkpoint
Only instruct models are listed: Hy4-preview and Hy4-preview-FP8. The card mentions no base checkpoint. Train from the unquantized instruct weights.
Tooling
The finetune guide offers three routes: DeepSpeed with the Hugging Face Trainer, LLaMA-Factory and ms-swift. Full fine-tuning and LoRA are both supported. The default LoRA setup is rank 64, alpha 128, dropout 0.05.
LoRA target modules
- DeepSpeed route:
q_proj,k_proj,v_proj,o_proj - LLaMA-Factory route:
q_a_proj,q_b_proj,kv_a_proj_with_mqa,kv_b_proj,o_proj - ms-swift: set through its config
Chat template and data
Data is JSONL with message lists. The model uses the hy_v4 chat template, and reasoning_effort selects high (slow thinking) or no_think (fast). The guide says the system prompt for training and inference is empty by default, though you can customize it.
MoE note
The router's e_score_correction_bias is a buffer that the training script loads automatically. The guide says not to ignore a failure to load it.
License
The card lists Apache License 2.0, so derivatives carry no extra limits beyond that license.
Memory to fine-tune Hy4, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Hy4-preview | 780.0B | 11622 GBNo live fit | 1453 GBNo live fit | 375 GBRTX A6000 × 8 · $2.90/hr | 4.7 × 10^21 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/tencent/Hy4-preview
- https://huggingface.co/tencent/Hy4-preview/raw/main/finetune/README.md
Updated 2026-10-07.