To fine-tune an LLM today you pick three things: a method (full fine-tuning, LoRA or QLoRA, then possibly preference tuning such as DPO or GRPO), a framework that implements it, and a GPU big enough for the memory that method needs. This guide maps the main open-source frameworks and methods, and puts the GPU memory numbers each project publishes side by side so you can size a run before you rent anything.
If you are new to the topic, read what is AI model fine-tuning first. Everything below is built from each project's own README or docs, cited in Sources, and "not stated" means the page we read did not say.
TL;DR
- Smallest GPU, fastest start: Unsloth for single-GPU LoRA and QLoRA. Its docs list a 7B model at 5 GB of VRAM with QLoRA (4-bit), as an absolute minimum.
- Config-file workflows with many methods: LLaMA-Factory and Axolotl (compared in Axolotl vs Unsloth vs torchtune) cover SFT, preference tuning and multi-GPU from YAML.
- Hugging Face ecosystem and RL methods: TRL provides SFT, DPO, GRPO and reward trainers on top of Accelerate, DeepSpeed and FSDP.
- Adapters under the hood: PEFT methods is the Hugging Face library most of the above use for LoRA and QLoRA.
- Models that do not fit on one GPU: DeepSpeed ZeRO or FSDP shard state across GPUs; Megatron-LM is for cluster-scale pretraining and large MoE.
- torchtune is no longer maintained, per its own README, so do not start a new project on it.
Step one: pick the method
The method sets the memory bill more than the framework does.
| Method | What it trains | Memory profile | Read more |
|---|---|---|---|
| Full fine-tuning | Every weight | Highest: weights, gradients and optimizer states | DeepSpeed guide |
| LoRA | Small adapter matrices, base weights frozen | Far lower than full | LoRA guide |
| QLoRA | LoRA on a 4-bit (or 8-bit) quantized base | Lowest | LoRA guide |
| Other PEFT methods | Adapters, prompt tuning, IA3 and more | Low | PEFT methods |
| DPO and other preference methods | Learns from chosen vs rejected answers | Similar to SFT plus a reference model | DPO vs PPO |
| GRPO | Reinforcement learning with group-relative rewards | Higher: generation plus training | GRPO explained |
The table is our summary. The glossary has entries for LoRA, QLoRA and quantization.
Step two: compare the frameworks
| Framework | Best for | Methods listed in its README | Multi-GPU | License |
|---|---|---|---|---|
| Unsloth | Fast, low-VRAM single-GPU runs | LoRA, QLoRA, full fine-tuning, pretraining, RL, GRPO, DPO, FP8 | "Multi GPU setups" listed as supported | Apache-2.0 core; some components such as the Studio UI are AGPL-3.0 |
| Axolotl | YAML-driven recipes across many methods | Full, LoRA, QLoRA, GPTQ, QAT, DPO/IPO/KTO/ORPO, GRPO/GDPO, reward modelling | FSDP2 and DeepSpeed; multi-node via Torchrun or Ray | Apache-2.0 |
| LLaMA-Factory | Web UI and CLI over many models and methods | Pretraining, SFT, reward modeling, PPO, DPO, KTO, ORPO, SimPO; full, freeze, LoRA, QLoRA, OFT | DeepSpeed optional; FSDP+QLoRA noted for 70B on 2x24GB | Apache-2.0 |
| TRL | Post-training with Hugging Face tools | SFT, DPO, GRPO, reward trainers, KTO via CLI | DDP, DeepSpeed ZeRO, FSDP via Accelerate | Apache-2.0 |
| PEFT | The adapter layer other tools build on | LoRA, QLoRA support, IA3, soft prompts, adapters | Works with Accelerate; DeepSpeed offload appears in its memory table | Apache-2.0 |
| DeepSpeed | Sharding state when a model does not fit | Not a trainer; ZeRO stages and offload | Core feature | Apache-2.0 |
| Megatron-LM | Cluster-scale pretraining, large MoE | Pretraining; SFT and LoRA via Megatron Bridge | TP, PP, DP, EP, CP | Apache |
| torchtune | Nothing new: unmaintained | Full, LoRA/QLoRA, DPO, PPO, GRPO, QAT | torchrun | BSD-3-Clause |
Sources for each cell are the repository READMEs listed below. Frameworks stack: TRL uses PEFT, and TRL, Axolotl and LLaMA-Factory all offer DeepSpeed and FSDP for multi-GPU. Unsloth is the one that ships its own optimized kernels, and its README claims 2x faster training with 70% less VRAM with no accuracy loss. Treat that as a vendor claim: the page gives no baseline hardware, so we have not verified it.
GPU memory: what each project publishes
Run the right numbers before you pick a card. These tables come from the projects themselves, and their conditions differ, so do not compare rows across tables as if they were one benchmark.
Unsloth: VRAM by model size
Unsloth's requirements page lists these as absolute minimums and notes some models need more:
| Model parameters | QLoRA (4-bit) VRAM | LoRA (16-bit) VRAM |
|---|---|---|
| 3B | 3.5 GB | 8 GB |
| 7B | 5 GB | 19 GB |
| 8B | 6 GB | 22 GB |
| 14B | 8.5 GB | 33 GB |
| 27B | 22 GB | 64 GB |
| 32B | 26 GB | 76 GB |
| 70B | 41 GB | 164 GB |
| 405B | 237 GB | 950 GB |
The page lists NVIDIA GPUs from 2018 onward with CUDA capability 7.0 or higher (including Blackwell RTX 50 and DGX Spark), and AMD and Intel GPUs through their own guides.
LLaMA-Factory: estimated memory by method
LLaMA-Factory's README marks this table as estimated:
| Method | Bits | 7B | 14B | 30B | 70B |
|---|---|---|---|---|---|
| Full (bf16 or fp16) | 32 | 120 GB | 240 GB | 600 GB | 1200 GB |
| Full (pure_bf16) | 16 | 60 GB | 120 GB | 300 GB | 600 GB |
| Freeze / LoRA / GaLore / APOLLO / BAdam / OFT | 16 | 16 GB | 32 GB | 64 GB | 160 GB |
| QLoRA | 8 | 10 GB | 20 GB | 40 GB | 80 GB |
| QLoRA | 4 | 6 GB | 12 GB | 24 GB | 48 GB |
| QLoRA | 2 | 4 GB | 8 GB | 16 GB | 24 GB |
Notice how the Unsloth and LLaMA-Factory numbers differ for the same model (a 7B LoRA is 19 GB in one table and 16 GB in the other). They come from different setups and assumptions, which is why you should treat any of them as a sizing starting point and test on your own sequence length and batch size.
torchtune: measured peak memory (Llama 3.1 8B, one RTX 4090)
The torchtune README published measured peak memory per GPU: 18.9 GiB for full fine-tuning, 16.2 GiB for LoRA and 7.4 GiB for QLoRA, on a single 4090. Another row shows Llama 3.1 70B LoRA on 8x A100 at 27.6 GiB per GPU. The project has stopped active development, so these are useful as reference points only.
PEFT: LoRA vs full on a 3B model
The PEFT README shows, on an A100 80GB with more than 64 GB of CPU RAM, T0_3B (3B parameters) using 47.14 GB of GPU memory for full fine-tuning, 14.4 GB with LoRA, and 9.8 GB with LoRA plus DeepSpeed CPU offloading (17.8 GB of CPU memory). It also shows that full fine-tuning of the 7B and 12B models in that table runs out of GPU memory on that card while LoRA fits at 32 GB and 56 GB.
Rule of thumb from the ZeRO paper (computed)
For full fine-tuning with Adam in mixed precision, the ZeRO paper counts 2 + 2 + 12 = 16 bytes per parameter for weights, gradients and optimizer states. Computed for a 7B model, that is about 112 GB before activations, which matches the order of magnitude of LLaMA-Factory's 120 GB row. Sharding across GPUs divides it, as covered in the DeepSpeed guide.
TRL, Axolotl, Megatron
TRL's and Axolotl's READMEs, as we read them, do not publish a memory table; both rely on PEFT and quantization for low-memory runs, and TRL's README states "training on large models with modest hardware via quantization and LoRA/QLoRA". Megatron-LM publishes throughput (up to 47% MFU on H100 clusters) rather than per-model memory. See the Megatron-LM guide.
Install and first run, from each project
# Unsloth (core, Linux/WSL)
uv pip install unsloth --torch-backend=auto
# Axolotl (README: Python 3.12 recommended)
uv pip install --no-build-isolation axolotl[deepspeed]
axolotl fetch examples
axolotl train examples/llama-3/lora-1b.yml
# LLaMA-Factory
git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
# TRL
pip install trl
trl sft --help
# PEFT
pip install peft
# DeepSpeed
pip install deepspeed
Axolotl's install line comes after a uv venv and a matching PyTorch install; follow the README for the full sequence. LLaMA-Factory also has a web UI via llamafactory-cli webui.
Which one should you use?
| Your situation | Pick | Why |
|---|---|---|
| One consumer or workstation GPU, 7B to 14B model | Unsloth with QLoRA | Lowest published VRAM floor (5 GB for 7B) |
| You want a UI or YAML and many model families | LLaMA-Factory or Axolotl | Configuration over code |
| You live in Hugging Face code and want DPO or GRPO | TRL (with PEFT) | Trainers and Accelerate integration |
| A 70B model on a few GPUs | LoRA or QLoRA with FSDP or DeepSpeed | LLaMA-Factory notes FSDP+QLoRA for 70B on 2x24GB |
| Full fine-tune that does not fit one GPU | DeepSpeed ZeRO-2 or ZeRO-3, or FSDP | Shards state across GPUs |
| Pretraining or very large MoE on a cluster | Megatron Core | Tensor, pipeline, context and expert parallelism |
Hardware sizing links: how much VRAM do I need for LLMs, the RTX 4090, L40S and H100 pages, and the VRAM calculator for consumer cards. For model weights you plan to adapt, browse /models.
Run it on a cloud GPU
QLoRA on a 7B to 14B model fits a 24 GB card by the tables above, while 70B-class LoRA wants an 80 GB card or several. Check what is available to rent now.
FAQ
How do I fine-tune an LLM?
Choose a method (usually LoRA or QLoRA), a framework (Unsloth, Axolotl, LLaMA-Factory or TRL), prepare an instruction or preference dataset, and run training on a GPU with enough memory. Then evaluate and merge or export the adapter.
How much GPU memory do I need to fine-tune a 7B model?
It depends on the method. Unsloth's docs list 5 GB for QLoRA (4-bit) and 19 GB for 16-bit LoRA as minimums for a 7B model. LLaMA-Factory's estimates for 7B are 6 GB (QLoRA 4-bit), 16 GB (LoRA) and 120 GB for full fine-tuning in bf16 or fp16.
Is Unsloth better than Axolotl or LLaMA-Factory?
They overlap. Unsloth's focus is speed and low memory on a single GPU, Axolotl and LLaMA-Factory are configuration-driven toolkits with broader method lists. See the comparison post for details.
Should I use torchtune?
Its README says it is no longer actively maintained and that development wound down in 2025, so choose a maintained framework for new work.
Do these frameworks support multiple GPUs?
Yes. Axolotl lists FSDP2 and DeepSpeed, TRL supports DDP, DeepSpeed ZeRO and FSDP through Accelerate, LLaMA-Factory has DeepSpeed as an optional dependency, and Unsloth lists multi-GPU setups as supported.
When do I need DeepSpeed or Megatron?
When the model states do not fit on one GPU. DeepSpeed ZeRO and FSDP shard them across GPUs. Megatron-LM is for cluster-scale pretraining and large MoE models, not typical LoRA runs.
Sources
- Unsloth README (license, hardware, methods, vendor claims): https://github.com/unslothai/unsloth
- Unsloth requirements (VRAM table, supported GPUs): https://unsloth.ai/docs/get-started/beginner-start-here/unsloth-requirements
- Axolotl README (methods, multi-GPU, install, train): https://github.com/axolotl-ai-cloud/axolotl
- LLaMA-Factory README (memory table, methods, license, commands): https://github.com/hiyouga/LLaMA-Factory
- TRL README (trainers, distributed support, PEFT): https://github.com/huggingface/trl
- PEFT README (methods, memory examples): https://github.com/huggingface/peft
- torchtune README (maintenance status, memory figures): https://github.com/pytorch/torchtune
- DeepSpeed README: https://github.com/deepspeedai/DeepSpeed
- ZeRO paper (16 bytes per parameter for mixed-precision Adam): https://arxiv.org/abs/1910.02054
- Megatron-LM README (MFU figure, parallelism): https://github.com/NVIDIA/Megatron-LM