Use Unsloth when you want the fastest, lowest-VRAM fine-tune on a single GPU. Use Axolotl when you want one YAML file to drive a repeatable, multi-GPU pipeline. Do not start a new project on torchtune: its own README says it is no longer actively maintained.
This comparison uses each project's own README and docs as of October 2026. For the full landscape of tools, see the LLM fine-tuning frameworks overview, and for the basics, what AI model fine-tuning is.
TL;DR
- Unsloth: single-GPU speed and low VRAM, notebooks and a desktop UI. Core is Apache-2.0, Studio is AGPL-3.0.
- Axolotl: one YAML for preprocessing, training, evaluation, quantization and inference, with multi-GPU and multi-node support and many preference and RL methods. Apache-2.0.
- torchtune: PyTorch-native recipes with clean, hackable code and published memory tables, but development wound down in 2025. Fine as a reference, risky as a base.
- Many teams use two of them: Unsloth to iterate on one card, Axolotl to scale the same idea to a node.
The status check first
torchtune's README opens with a warning that it "is no longer actively maintained" and says development wound down in 2025. The linked announcement, an issue titled "The future of torchtune" opened on July 15, 2025, states that active development stopped effective immediately, that critical bug fixes and security patches would continue during 2025, that no new features would be added, and that a new product in a new repo would carry the work forward. The README's model list tops out at Llama 4 Scout, Qwen3 and Phi4, and its recent updates list ends in May 2025 with Qwen3 support. Anything newer than that is not there.
Axolotl and Unsloth are both actively updated. The Axolotl README's latest update note is dated October 2026 (distributed training improvements such as Ulysses sequence parallelism). The Unsloth README lists support for current model families such as Gemma 4, DeepSeek-V4 and Qwen3.x.
Side-by-side
| Unsloth | Axolotl | torchtune | |
|---|---|---|---|
| Status (Oct 2026) | Active | Active | No longer actively maintained |
| License | Core Apache-2.0, Studio AGPL-3.0 | Apache-2.0 | See repo |
| Interface | Python notebooks, Studio UI | YAML config plus CLI | tune run recipes plus YAML |
| Methods | LoRA, QLoRA, full, pretraining, GRPO, DPO, FP8 | Full, LoRA, QLoRA, QAT, FP8, DPO, IPO, KTO, ORPO, GRPO, reward models | Full, LoRA, QLoRA, DPO, QAT, knowledge distillation |
| Multi-GPU | Supported per README | Core feature, including multi-node | Distributed recipes (8x A100 in its tables) |
| Multimodal | Diffusion, TTS, embedding, audio | Vision-language and audio | Llama 3.2 Vision |
| Hardware | NVIDIA, AMD, Intel, Apple, CPU | NVIDIA Ampere or newer, or AMD | NVIDIA, with an XPU example |
Sources for each cell are the three READMEs listed at the end.
Unsloth
Unsloth optimizes the hot paths of LoRA and QLoRA training so a given fine-tune fits on smaller GPUs. The project's own VRAM table says QLoRA on a 7B model needs about 5 GB and LoRA about 19 GB, and its headline claim is training "2x faster with 70% less VRAM". Those are the project's numbers on its notebooks. The details, install commands and the full table are in the Unsloth guide.
Choose it when you are on one GPU, want to be training in minutes, or want a UI. Be careful if you plan to distribute Studio, because it carries AGPL-3.0 terms.
Axolotl
Axolotl is configuration first. A single YAML file covers dataset preprocessing, training, evaluation, quantization and inference. The install in the README uses uv:
curl -LsSf https://astral.sh/uv/install.sh | sh
export UV_TORCH_BACKEND=cu130
uv venv --python 3.12
source .venv/bin/activate
uv pip install torch==2.14.0 torchvision
uv pip install --no-build-isolation axolotl[deepspeed]
There is also a Docker image:
docker run --gpus '"all"' --ipc=host --rm -it axolotlai/axolotl:main-latest
The quickstart from the README:
axolotl fetch examples
axolotl fetch deepspeed_configs # optional
axolotl train examples/llama-3/lora-1b.yml
The Axolotl docs show a small config of this shape (a partial example from their getting-started page):
base_model: NousResearch/Llama-3.2-1B
load_in_8bit: true
adapter: lora
datasets:
- path: teknium/GPT4-LLM-Cleaned
type: alpaca
dataset_prepared_path: last_run_prepared
val_set_size: 0.1
output_dir: ./outputs/lora-out
The same file then drives axolotl preprocess, axolotl inference and axolotl merge-lora. Requirements per the README: an NVIDIA Ampere or newer GPU (for bf16 and Flash Attention) or an AMD GPU, Python 3.12 or newer, and a recent PyTorch (the README recommends 2.13.0 while its install command pins 2.14.0, so follow the install block).
Its method list is the widest of the three: full fine-tuning, LoRA, QLoRA, GPTQ, QAT, FP8 mixed precision and MoE LoRA; DPO, IPO, KTO and ORPO for preference tuning; GRPO and GDPO for RL; and reward modelling. Performance features include multipacking, several Flash Attention versions, Liger Kernel, sequence parallelism, and multi-GPU and multi-node training. See Flash Attention 2 vs 3 for the attention backends and the DeepSpeed guide for the sharding layer it builds on.
Choose it when the run must be reproducible from a committed file, when you scale past one GPU, or when you need a preference or RL method beyond SFT.
torchtune
torchtune is a PyTorch-native library of training recipes. Commands from its README:
pip install torch torchvision torchao
pip install torchtune
tune run lora_finetune_single_device --config llama3_1/8B_lora_single_device
tune run --nproc_per_node 2 full_finetune_distributed --config llama3_1/8B_full
tune run lora_dpo_single_device --config llama3_1/8B_dpo_single_device
Its main lasting value is the published memory and throughput tables, which are the project's own measurements (Llama 3.1, batch size 2, packed dataset at sequence length 2048, torch compile on):
| Model | Method | Hardware | Peak memory per GPU | Tokens/sec |
|---|---|---|---|---|
| Llama 3.1 8B | Full fine-tune | 1x RTX 4090 | 18.9 GiB | 1650 |
| Llama 3.1 8B | LoRA | 1x RTX 4090 | 16.2 GiB | 3083 |
| Llama 3.1 8B | QLoRA | 1x RTX 4090 | 7.4 GiB | 2413 |
| Llama 3.1 70B | LoRA | 8x A100 | 27.6 GiB | 3497 |
| Llama 3.1 405B | QLoRA | 8x A100 | 44.8 GB | 653 |
These come from torchtune's own README, were measured on its recipes, and are not comparable one-to-one with Unsloth's or Axolotl's numbers. The README also shows a Llama 3.2 3B ladder on an A100 where stacking packing, compile, activation checkpointing and other flags moves peak memory from 25.5 GiB for the baseline to 4.6 GiB for QLoRA. The point worth taking from it: most of the memory savings come from standard techniques you can also get in TRL or Axolotl.
Choose it only if you want readable PyTorch recipes to study or fork, and accept that you own maintenance.
Which one should you pick
- One GPU, first fine-tune, fastest path: Unsloth.
- A config you can commit, rerun and hand to a teammate: Axolotl.
- Multi-GPU or multi-node SFT, DPO or GRPO: Axolotl, or TRL with DeepSpeed or FSDP (see the TRL guide).
- Many model families and a web UI: consider LLaMA-Factory.
- Studying how a recipe works in plain PyTorch: torchtune, as a reference.
For the techniques behind all three, read the glossary entries on LoRA, QLoRA and FSDP.
Run it on a cloud GPU
Aquanode manages and optimizes GPUs for training and inference workloads. A 24 GB card is enough to try any of these tools on a 7B to 8B model with QLoRA, and an 80 GB card lets you move on to 16-bit LoRA at larger sizes. Live availability and pricing are below.
FAQ
Is torchtune dead?
Its README says it is no longer actively maintained, and the July 2025 announcement says no new features will be added. The code still works for the models it supports, but expect no support for newer model families.
Axolotl or Unsloth for a first fine-tune?
Unsloth gets you to a running job faster on one GPU. Axolotl is the better choice once you want a versioned YAML and multiple GPUs. Many people start in one and move to the other.
Can I use Unsloth and Axolotl together?
Axolotl lists Liger Kernel and other kernel-level options of its own. Check the current Axolotl docs for whether a given Unsloth optimization is supported in your version before combining them.
Do they all support LoRA and QLoRA?
Yes. Unsloth, Axolotl and torchtune all list LoRA and QLoRA, per their READMEs.
Which supports GRPO?
Unsloth and Axolotl list GRPO in their READMEs. torchtune's README does not list it. TRL also provides a GRPOTrainer, covered in the TRL guide.
Sources
- Axolotl GitHub README: https://github.com/axolotl-ai-cloud/axolotl
- Axolotl getting started docs: https://docs.axolotl.ai/docs/getting-started.html
- Unsloth GitHub README: https://github.com/unslothai/unsloth
- Unsloth requirements and VRAM table: https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements
- torchtune GitHub README (maintenance notice, tables, commands): https://github.com/pytorch/torchtune
- torchtune announcement, "The future of torchtune": https://github.com/pytorch/torchtune/issues/2883