LLaMA-Factory is an open-source framework that fine-tunes more than a hundred LLM and vision-language models through one YAML-and-CLI workflow or a browser UI called LLaMA Board. You pick a model, a method (SFT, DPO, PPO and others), a dataset and a precision, and it handles the training loop.
This guide covers the install, the commands, the methods, and the GPU memory table the project publishes. For the wider tool landscape, see the LLM fine-tuning frameworks overview. For the concepts, see what AI model fine-tuning is.
TL;DR
- LLaMA-Factory is the broadest-coverage option: pre-training, SFT, reward modeling, PPO, DPO, KTO, ORPO and SimPO, each with full, freeze, LoRA and QLoRA variants (project README).
- You can run it with no code:
llamafactory-cli train my_config.yamlorllamafactory-cli webui. - The project's estimate for 16-bit LoRA is about 2 GB of VRAM per billion parameters, and for 4-bit QLoRA about half a GB per billion (project table, labelled "estimated").
- Apache-2.0 licensed. Model weights keep their own licenses.
- Pick it for variety and a UI. Pick Unsloth for single-GPU speed, Axolotl for config-first multi-GPU pipelines, TRL for code-level control.
What is LLaMA-Factory
The repository (hiyouga/LLaMA-Factory) is a unified fine-tuning toolkit. Its README lists support for model families including LLaMA, Mistral, Mixtral, Qwen, DeepSeek, Gemma, GLM and Phi, plus multimodal models such as Qwen-VL, LLaVA and InternVL. It handles multi-turn dialogue, tool use, and image, video and audio understanding.
Training algorithms and tricks named in the README include GaLore, BAdam, APOLLO, Muon, DoRA, LoRA+ and PiSSA, with FlashAttention-2 and Liger Kernel as speed options. Monitoring integrations include TensorBoard, W&B, MLflow and SwanLab. For inference it exposes an OpenAI-style API backed by vLLM or SGLang, which connects to serving LLMs with vLLM and the SGLang guide.
The newest changelog entry in the README (as of this writing) announces a Megatron-core training backend through an adapter. That is aimed at large-scale runs, see the Megatron-LM guide.
Supported methods
From the README's table, every one of these approaches supports full-tuning, freeze-tuning, LoRA, QLoRA, OFT and QOFT:
- Pre-training
- Supervised fine-tuning (SFT)
- Reward modeling
- PPO
- DPO
- KTO
- ORPO
- SimPO
For the difference between LoRA and QLoRA, see the glossary entries on LoRA and QLoRA. For when to use DPO instead of PPO, see DPO vs PPO.
Install LLaMA-Factory
From the project README, install from source:
git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
pip install -r requirements/metrics.txt
Optional extras are metrics and deepspeed.
The README also gives a Docker command:
docker run -it --rm --gpus=all --ipc=host hiyouga/llamafactory:latest
The image described in the README is built on Ubuntu 22.04 with CUDA 12.4, Python 3.11, PyTorch 2.6.0 and Flash-attn 2.7.4. If you need a newer CUDA or PyTorch for a Blackwell GPU, a source install in your own environment is the safer route. Check the README for the current image tags.
Run it: CLI and web UI
The README quickstart uses a Qwen3 LoRA example:
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml
llamafactory-cli webui
The first command trains a LoRA adapter, the second chats with it, the third merges the adapter into the base weights, and the last opens the Gradio-based LLaMA Board.
The training YAML in the repository for that example is a LoRA SFT config for Qwen3-4B-Instruct-2507. As summarized from the file: LoRA rank 8 on all target modules, the identity and alpaca_en_demo datasets capped at 1,000 samples, the qwen3_nothink template with a 2,048-token cutoff, batch size 1 with 8 gradient accumulation steps, learning rate 1.0e-4, 3 epochs, cosine scheduler with 10 percent warmup, bf16, output to saves/qwen3-4b/lora/sft. Open the file in the repo to copy the exact keys before editing, because option names change between releases.
A normal flow looks like this:
- Register your dataset in the project's dataset info file, in a format it supports (Alpaca-style instruction records or ShareGPT-style conversations are described in the README's data docs).
- Copy an example YAML, change the model path, dataset name and template.
- Run
llamafactory-cli trainand watch the loss. - Chat with the adapter, then export a merged model.
- Serve it through the OpenAI-style API with vLLM or SGLang.
On a remote machine, run the web UI and tunnel the port rather than exposing it publicly.
How much GPU memory does LLaMA-Factory need
This is the project's own table from the README. The README labels the figures as estimates, and gives a formula where x is the parameter count in billions.
| Method | Bits | 7B | 14B | 30B | 70B | Formula |
|---|---|---|---|---|---|---|
| Full (bf16 or fp16) | 32 | 120 GB | 240 GB | 600 GB | 1200 GB | 18x GB |
| Full (pure_bf16) | 16 | 60 GB | 120 GB | 300 GB | 600 GB | 8x GB |
| Freeze, LoRA, GaLore, APOLLO, BAdam, OFT | 16 | 16 GB | 32 GB | 64 GB | 160 GB | 2x GB |
| QLoRA or QOFT | 8 | 10 GB | 20 GB | 40 GB | 80 GB | x GB |
| QLoRA or QOFT | 4 | 6 GB | 12 GB | 24 GB | 48 GB | x/2 GB |
| QLoRA or QOFT | 2 | 4 GB | 8 GB | 16 GB | 24 GB | x/4 GB |
What that means for GPU choice (computed from the table above, not measured by us):
- 7B with 4-bit QLoRA, 6 GB: fits on a 24 GB card such as the RTX 4090 with generous room for sequence length and batch size.
- 7B with 16-bit LoRA, 16 GB: also fits on 24 GB.
- 14B with 16-bit LoRA, 32 GB: needs a 48 GB card such as the L40S.
- 70B with 4-bit QLoRA, 48 GB: sits at the edge of a 48 GB card, so an 80 GB H100 is the comfortable choice.
- 7B full fine-tuning in pure bf16, 60 GB: needs an 80 GB card.
These numbers are estimates and do not include long contexts, large batches or the activation memory of vision inputs. Use the VRAM calculator for the RTX 4090 or read how much VRAM you need for LLMs to cross-check. Background: VRAM and quantization.
LLaMA-Factory vs Unsloth, Axolotl and TRL
| Tool | Strength | Interface |
|---|---|---|
| LLaMA-Factory | Widest model and method coverage, built-in web UI | YAML, CLI, LLaMA Board |
| Unsloth | Single-GPU speed and low VRAM | Notebooks, Studio |
| Axolotl | Config-first, multi-GPU and multi-node | YAML, CLI |
| TRL | Code-level trainers, the base layer others build on | Python |
LLaMA-Factory has no published speed benchmark in the README that we can cite for a head-to-head, so we make none. If raw throughput on one GPU is the priority, compare with the Unsloth guide and measure both on your model. The three-way view, including torchtune's status, is in Axolotl vs Unsloth vs torchtune, and the TRL route is in the TRL guide.
Pick LLaMA-Factory when you want to try many models or methods quickly, when teammates who do not write Python need to launch runs from a UI, or when you want one tool that goes from SFT to DPO to export without changing stacks.
Run it on a cloud GPU
Aquanode manages and optimizes GPUs for training and inference workloads. Match the card to the table above: 24 GB for 7B QLoRA or LoRA, 48 GB for 14B LoRA, 80 GB for 70B QLoRA. Live availability and pricing are below.
FAQ
Is LLaMA-Factory free?
Yes. The code is Apache-2.0 per the README. Model weights are governed by their own licenses, so check the license of the base model you fine-tune.
Can I use LLaMA-Factory without writing code?
Yes. The llamafactory-cli webui command opens LLaMA Board, a Gradio interface for configuring and launching runs, and the CLI takes a YAML file.
How much VRAM do I need to fine-tune a 7B model?
The project's estimate is about 6 GB for 4-bit QLoRA and about 16 GB for 16-bit LoRA, before long contexts or large batches.
Does it support DPO and GRPO?
DPO, KTO, ORPO, SimPO, PPO and reward modeling are listed in the README. GRPO is not in the method list we read, so for GRPO see the TRL guide or GRPO explained.
Can it train across several GPUs?
The README lists a DeepSpeed extra and a Megatron-core backend. Consult the project docs for the launch commands for your setup, see also the DeepSpeed guide.
Sources
- LLaMA-Factory GitHub README (install, commands, methods, memory table, license): https://github.com/hiyouga/LLaMA-Factory
- LLaMA-Factory example config, Qwen3 LoRA SFT: https://github.com/hiyouga/LLaMA-Factory/blob/main/examples/train_lora/qwen3_lora_sft.yaml