LLM Fine-Tuning Frameworks Compared (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

To fine-tune an LLM today you pick three things: a method (full fine-tuning, LoRA or QLoRA, then possibly preference tuning such as DPO or GRPO), a framework that implements it, and a GPU big enough for the memory that method needs. This guide maps the main open-source frameworks and methods, and puts the GPU memory numbers each project publishes side by side so you can size a run before you rent anything.

If you are new to the topic, read what is AI model fine-tuning first. Everything below is built from each project's own README or docs, cited in Sources, and "not stated" means the page we read did not say.

TL;DR

  • Smallest GPU, fastest start: Unsloth for single-GPU LoRA and QLoRA. Its docs list a 7B model at 5 GB of VRAM with QLoRA (4-bit), as an absolute minimum.
  • Config-file workflows with many methods: LLaMA-Factory and Axolotl (compared in Axolotl vs Unsloth vs torchtune) cover SFT, preference tuning and multi-GPU from YAML.
  • Hugging Face ecosystem and RL methods: TRL provides SFT, DPO, GRPO and reward trainers on top of Accelerate, DeepSpeed and FSDP.
  • Adapters under the hood: PEFT methods is the Hugging Face library most of the above use for LoRA and QLoRA.
  • Models that do not fit on one GPU: DeepSpeed ZeRO or FSDP shard state across GPUs; Megatron-LM is for cluster-scale pretraining and large MoE.
  • torchtune is no longer maintained, per its own README, so do not start a new project on it.

Step one: pick the method

The method sets the memory bill more than the framework does.

MethodWhat it trainsMemory profileRead more
Full fine-tuningEvery weightHighest: weights, gradients and optimizer statesDeepSpeed guide
LoRASmall adapter matrices, base weights frozenFar lower than fullLoRA guide
QLoRALoRA on a 4-bit (or 8-bit) quantized baseLowestLoRA guide
Other PEFT methodsAdapters, prompt tuning, IA3 and moreLowPEFT methods
DPO and other preference methodsLearns from chosen vs rejected answersSimilar to SFT plus a reference modelDPO vs PPO
GRPOReinforcement learning with group-relative rewardsHigher: generation plus trainingGRPO explained

The table is our summary. The glossary has entries for LoRA, QLoRA and quantization.

Step two: compare the frameworks

FrameworkBest forMethods listed in its READMEMulti-GPULicense
UnslothFast, low-VRAM single-GPU runsLoRA, QLoRA, full fine-tuning, pretraining, RL, GRPO, DPO, FP8"Multi GPU setups" listed as supportedApache-2.0 core; some components such as the Studio UI are AGPL-3.0
AxolotlYAML-driven recipes across many methodsFull, LoRA, QLoRA, GPTQ, QAT, DPO/IPO/KTO/ORPO, GRPO/GDPO, reward modellingFSDP2 and DeepSpeed; multi-node via Torchrun or RayApache-2.0
LLaMA-FactoryWeb UI and CLI over many models and methodsPretraining, SFT, reward modeling, PPO, DPO, KTO, ORPO, SimPO; full, freeze, LoRA, QLoRA, OFTDeepSpeed optional; FSDP+QLoRA noted for 70B on 2x24GBApache-2.0
TRLPost-training with Hugging Face toolsSFT, DPO, GRPO, reward trainers, KTO via CLIDDP, DeepSpeed ZeRO, FSDP via AccelerateApache-2.0
PEFTThe adapter layer other tools build onLoRA, QLoRA support, IA3, soft prompts, adaptersWorks with Accelerate; DeepSpeed offload appears in its memory tableApache-2.0
DeepSpeedSharding state when a model does not fitNot a trainer; ZeRO stages and offloadCore featureApache-2.0
Megatron-LMCluster-scale pretraining, large MoEPretraining; SFT and LoRA via Megatron BridgeTP, PP, DP, EP, CPApache
torchtuneNothing new: unmaintainedFull, LoRA/QLoRA, DPO, PPO, GRPO, QATtorchrunBSD-3-Clause

Sources for each cell are the repository READMEs listed below. Frameworks stack: TRL uses PEFT, and TRL, Axolotl and LLaMA-Factory all offer DeepSpeed and FSDP for multi-GPU. Unsloth is the one that ships its own optimized kernels, and its README claims 2x faster training with 70% less VRAM with no accuracy loss. Treat that as a vendor claim: the page gives no baseline hardware, so we have not verified it.

GPU memory: what each project publishes

Run the right numbers before you pick a card. These tables come from the projects themselves, and their conditions differ, so do not compare rows across tables as if they were one benchmark.

Unsloth: VRAM by model size

Unsloth's requirements page lists these as absolute minimums and notes some models need more:

Model parametersQLoRA (4-bit) VRAMLoRA (16-bit) VRAM
3B3.5 GB8 GB
7B5 GB19 GB
8B6 GB22 GB
14B8.5 GB33 GB
27B22 GB64 GB
32B26 GB76 GB
70B41 GB164 GB
405B237 GB950 GB

The page lists NVIDIA GPUs from 2018 onward with CUDA capability 7.0 or higher (including Blackwell RTX 50 and DGX Spark), and AMD and Intel GPUs through their own guides.

LLaMA-Factory: estimated memory by method

LLaMA-Factory's README marks this table as estimated:

MethodBits7B14B30B70B
Full (bf16 or fp16)32120 GB240 GB600 GB1200 GB
Full (pure_bf16)1660 GB120 GB300 GB600 GB
Freeze / LoRA / GaLore / APOLLO / BAdam / OFT1616 GB32 GB64 GB160 GB
QLoRA810 GB20 GB40 GB80 GB
QLoRA46 GB12 GB24 GB48 GB
QLoRA24 GB8 GB16 GB24 GB

Notice how the Unsloth and LLaMA-Factory numbers differ for the same model (a 7B LoRA is 19 GB in one table and 16 GB in the other). They come from different setups and assumptions, which is why you should treat any of them as a sizing starting point and test on your own sequence length and batch size.

torchtune: measured peak memory (Llama 3.1 8B, one RTX 4090)

The torchtune README published measured peak memory per GPU: 18.9 GiB for full fine-tuning, 16.2 GiB for LoRA and 7.4 GiB for QLoRA, on a single 4090. Another row shows Llama 3.1 70B LoRA on 8x A100 at 27.6 GiB per GPU. The project has stopped active development, so these are useful as reference points only.

PEFT: LoRA vs full on a 3B model

The PEFT README shows, on an A100 80GB with more than 64 GB of CPU RAM, T0_3B (3B parameters) using 47.14 GB of GPU memory for full fine-tuning, 14.4 GB with LoRA, and 9.8 GB with LoRA plus DeepSpeed CPU offloading (17.8 GB of CPU memory). It also shows that full fine-tuning of the 7B and 12B models in that table runs out of GPU memory on that card while LoRA fits at 32 GB and 56 GB.

Rule of thumb from the ZeRO paper (computed)

For full fine-tuning with Adam in mixed precision, the ZeRO paper counts 2 + 2 + 12 = 16 bytes per parameter for weights, gradients and optimizer states. Computed for a 7B model, that is about 112 GB before activations, which matches the order of magnitude of LLaMA-Factory's 120 GB row. Sharding across GPUs divides it, as covered in the DeepSpeed guide.

TRL, Axolotl, Megatron

TRL's and Axolotl's READMEs, as we read them, do not publish a memory table; both rely on PEFT and quantization for low-memory runs, and TRL's README states "training on large models with modest hardware via quantization and LoRA/QLoRA". Megatron-LM publishes throughput (up to 47% MFU on H100 clusters) rather than per-model memory. See the Megatron-LM guide.

Install and first run, from each project

# Unsloth (core, Linux/WSL)
uv pip install unsloth --torch-backend=auto

# Axolotl (README: Python 3.12 recommended)
uv pip install --no-build-isolation axolotl[deepspeed]
axolotl fetch examples
axolotl train examples/llama-3/lora-1b.yml

# LLaMA-Factory
git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml

# TRL
pip install trl
trl sft --help

# PEFT
pip install peft

# DeepSpeed
pip install deepspeed

Axolotl's install line comes after a uv venv and a matching PyTorch install; follow the README for the full sequence. LLaMA-Factory also has a web UI via llamafactory-cli webui.

Which one should you use?

Your situationPickWhy
One consumer or workstation GPU, 7B to 14B modelUnsloth with QLoRALowest published VRAM floor (5 GB for 7B)
You want a UI or YAML and many model familiesLLaMA-Factory or AxolotlConfiguration over code
You live in Hugging Face code and want DPO or GRPOTRL (with PEFT)Trainers and Accelerate integration
A 70B model on a few GPUsLoRA or QLoRA with FSDP or DeepSpeedLLaMA-Factory notes FSDP+QLoRA for 70B on 2x24GB
Full fine-tune that does not fit one GPUDeepSpeed ZeRO-2 or ZeRO-3, or FSDPShards state across GPUs
Pretraining or very large MoE on a clusterMegatron CoreTensor, pipeline, context and expert parallelism

Hardware sizing links: how much VRAM do I need for LLMs, the RTX 4090, L40S and H100 pages, and the VRAM calculator for consumer cards. For model weights you plan to adapt, browse /models.

Run it on a cloud GPU

QLoRA on a 7B to 14B model fits a 24 GB card by the tables above, while 70B-class LoRA wants an 80 GB card or several. Check what is available to rent now.

FAQ

How do I fine-tune an LLM?

Choose a method (usually LoRA or QLoRA), a framework (Unsloth, Axolotl, LLaMA-Factory or TRL), prepare an instruction or preference dataset, and run training on a GPU with enough memory. Then evaluate and merge or export the adapter.

How much GPU memory do I need to fine-tune a 7B model?

It depends on the method. Unsloth's docs list 5 GB for QLoRA (4-bit) and 19 GB for 16-bit LoRA as minimums for a 7B model. LLaMA-Factory's estimates for 7B are 6 GB (QLoRA 4-bit), 16 GB (LoRA) and 120 GB for full fine-tuning in bf16 or fp16.

Is Unsloth better than Axolotl or LLaMA-Factory?

They overlap. Unsloth's focus is speed and low memory on a single GPU, Axolotl and LLaMA-Factory are configuration-driven toolkits with broader method lists. See the comparison post for details.

Should I use torchtune?

Its README says it is no longer actively maintained and that development wound down in 2025, so choose a maintained framework for new work.

Do these frameworks support multiple GPUs?

Yes. Axolotl lists FSDP2 and DeepSpeed, TRL supports DDP, DeepSpeed ZeRO and FSDP through Accelerate, LLaMA-Factory has DeepSpeed as an optional dependency, and Unsloth lists multi-GPU setups as supported.

When do I need DeepSpeed or Megatron?

When the model states do not fit on one GPU. DeepSpeed ZeRO and FSDP shard them across GPUs. Megatron-LM is for cluster-scale pretraining and large MoE models, not typical LoRA runs.

Sources

#fine-tuning#llm training#unsloth#axolotl#trl#llama-factory#deepspeed#peft

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.