LLaMA-Factory Guide: Fine-Tune 100+ LLMs (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

LLaMA-Factory is an open-source framework that fine-tunes more than a hundred LLM and vision-language models through one YAML-and-CLI workflow or a browser UI called LLaMA Board. You pick a model, a method (SFT, DPO, PPO and others), a dataset and a precision, and it handles the training loop.

This guide covers the install, the commands, the methods, and the GPU memory table the project publishes. For the wider tool landscape, see the LLM fine-tuning frameworks overview. For the concepts, see what AI model fine-tuning is.

TL;DR

  • LLaMA-Factory is the broadest-coverage option: pre-training, SFT, reward modeling, PPO, DPO, KTO, ORPO and SimPO, each with full, freeze, LoRA and QLoRA variants (project README).
  • You can run it with no code: llamafactory-cli train my_config.yaml or llamafactory-cli webui.
  • The project's estimate for 16-bit LoRA is about 2 GB of VRAM per billion parameters, and for 4-bit QLoRA about half a GB per billion (project table, labelled "estimated").
  • Apache-2.0 licensed. Model weights keep their own licenses.
  • Pick it for variety and a UI. Pick Unsloth for single-GPU speed, Axolotl for config-first multi-GPU pipelines, TRL for code-level control.

What is LLaMA-Factory

The repository (hiyouga/LLaMA-Factory) is a unified fine-tuning toolkit. Its README lists support for model families including LLaMA, Mistral, Mixtral, Qwen, DeepSeek, Gemma, GLM and Phi, plus multimodal models such as Qwen-VL, LLaVA and InternVL. It handles multi-turn dialogue, tool use, and image, video and audio understanding.

Training algorithms and tricks named in the README include GaLore, BAdam, APOLLO, Muon, DoRA, LoRA+ and PiSSA, with FlashAttention-2 and Liger Kernel as speed options. Monitoring integrations include TensorBoard, W&B, MLflow and SwanLab. For inference it exposes an OpenAI-style API backed by vLLM or SGLang, which connects to serving LLMs with vLLM and the SGLang guide.

The newest changelog entry in the README (as of this writing) announces a Megatron-core training backend through an adapter. That is aimed at large-scale runs, see the Megatron-LM guide.

Supported methods

From the README's table, every one of these approaches supports full-tuning, freeze-tuning, LoRA, QLoRA, OFT and QOFT:

  • Pre-training
  • Supervised fine-tuning (SFT)
  • Reward modeling
  • PPO
  • DPO
  • KTO
  • ORPO
  • SimPO

For the difference between LoRA and QLoRA, see the glossary entries on LoRA and QLoRA. For when to use DPO instead of PPO, see DPO vs PPO.

Install LLaMA-Factory

From the project README, install from source:

git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
pip install -r requirements/metrics.txt

Optional extras are metrics and deepspeed.

The README also gives a Docker command:

docker run -it --rm --gpus=all --ipc=host hiyouga/llamafactory:latest

The image described in the README is built on Ubuntu 22.04 with CUDA 12.4, Python 3.11, PyTorch 2.6.0 and Flash-attn 2.7.4. If you need a newer CUDA or PyTorch for a Blackwell GPU, a source install in your own environment is the safer route. Check the README for the current image tags.

Run it: CLI and web UI

The README quickstart uses a Qwen3 LoRA example:

llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml
llamafactory-cli webui

The first command trains a LoRA adapter, the second chats with it, the third merges the adapter into the base weights, and the last opens the Gradio-based LLaMA Board.

The training YAML in the repository for that example is a LoRA SFT config for Qwen3-4B-Instruct-2507. As summarized from the file: LoRA rank 8 on all target modules, the identity and alpaca_en_demo datasets capped at 1,000 samples, the qwen3_nothink template with a 2,048-token cutoff, batch size 1 with 8 gradient accumulation steps, learning rate 1.0e-4, 3 epochs, cosine scheduler with 10 percent warmup, bf16, output to saves/qwen3-4b/lora/sft. Open the file in the repo to copy the exact keys before editing, because option names change between releases.

A normal flow looks like this:

  1. Register your dataset in the project's dataset info file, in a format it supports (Alpaca-style instruction records or ShareGPT-style conversations are described in the README's data docs).
  2. Copy an example YAML, change the model path, dataset name and template.
  3. Run llamafactory-cli train and watch the loss.
  4. Chat with the adapter, then export a merged model.
  5. Serve it through the OpenAI-style API with vLLM or SGLang.

On a remote machine, run the web UI and tunnel the port rather than exposing it publicly.

How much GPU memory does LLaMA-Factory need

This is the project's own table from the README. The README labels the figures as estimates, and gives a formula where x is the parameter count in billions.

MethodBits7B14B30B70BFormula
Full (bf16 or fp16)32120 GB240 GB600 GB1200 GB18x GB
Full (pure_bf16)1660 GB120 GB300 GB600 GB8x GB
Freeze, LoRA, GaLore, APOLLO, BAdam, OFT1616 GB32 GB64 GB160 GB2x GB
QLoRA or QOFT810 GB20 GB40 GB80 GBx GB
QLoRA or QOFT46 GB12 GB24 GB48 GBx/2 GB
QLoRA or QOFT24 GB8 GB16 GB24 GBx/4 GB

What that means for GPU choice (computed from the table above, not measured by us):

  • 7B with 4-bit QLoRA, 6 GB: fits on a 24 GB card such as the RTX 4090 with generous room for sequence length and batch size.
  • 7B with 16-bit LoRA, 16 GB: also fits on 24 GB.
  • 14B with 16-bit LoRA, 32 GB: needs a 48 GB card such as the L40S.
  • 70B with 4-bit QLoRA, 48 GB: sits at the edge of a 48 GB card, so an 80 GB H100 is the comfortable choice.
  • 7B full fine-tuning in pure bf16, 60 GB: needs an 80 GB card.

These numbers are estimates and do not include long contexts, large batches or the activation memory of vision inputs. Use the VRAM calculator for the RTX 4090 or read how much VRAM you need for LLMs to cross-check. Background: VRAM and quantization.

LLaMA-Factory vs Unsloth, Axolotl and TRL

ToolStrengthInterface
LLaMA-FactoryWidest model and method coverage, built-in web UIYAML, CLI, LLaMA Board
UnslothSingle-GPU speed and low VRAMNotebooks, Studio
AxolotlConfig-first, multi-GPU and multi-nodeYAML, CLI
TRLCode-level trainers, the base layer others build onPython

LLaMA-Factory has no published speed benchmark in the README that we can cite for a head-to-head, so we make none. If raw throughput on one GPU is the priority, compare with the Unsloth guide and measure both on your model. The three-way view, including torchtune's status, is in Axolotl vs Unsloth vs torchtune, and the TRL route is in the TRL guide.

Pick LLaMA-Factory when you want to try many models or methods quickly, when teammates who do not write Python need to launch runs from a UI, or when you want one tool that goes from SFT to DPO to export without changing stacks.

Run it on a cloud GPU

Aquanode manages and optimizes GPUs for training and inference workloads. Match the card to the table above: 24 GB for 7B QLoRA or LoRA, 48 GB for 14B LoRA, 80 GB for 70B QLoRA. Live availability and pricing are below.

FAQ

Is LLaMA-Factory free?

Yes. The code is Apache-2.0 per the README. Model weights are governed by their own licenses, so check the license of the base model you fine-tune.

Can I use LLaMA-Factory without writing code?

Yes. The llamafactory-cli webui command opens LLaMA Board, a Gradio interface for configuring and launching runs, and the CLI takes a YAML file.

How much VRAM do I need to fine-tune a 7B model?

The project's estimate is about 6 GB for 4-bit QLoRA and about 16 GB for 16-bit LoRA, before long contexts or large batches.

Does it support DPO and GRPO?

DPO, KTO, ORPO, SimPO, PPO and reward modeling are listed in the README. GRPO is not in the method list we read, so for GRPO see the TRL guide or GRPO explained.

Can it train across several GPUs?

The README lists a DeepSpeed extra and a Megatron-core backend. Consult the project docs for the launch commands for your setup, see also the DeepSpeed guide.

Sources

#fine-tuning#llm training#llama-factory#lora#qlora#dpo

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.