Unsloth is an open-source library and desktop app for fine-tuning and running LLMs. Its pitch is lower memory use and faster training for LoRA, QLoRA, full fine-tuning and reinforcement learning, on a single GPU that you can rent by the hour.
This guide covers what Unsloth is, the install paths its own docs give, the VRAM table the project publishes, and where it fits next to the other fine-tuning stacks. For the wider map of tools, start at the LLM fine-tuning frameworks overview. If you want the concepts first, read what AI model fine-tuning is.
TL;DR
- Unsloth is the lowest-friction way to fine-tune a 3B to 32B model on one GPU. The project says QLoRA on a 7B model needs about 5 GB of VRAM and LoRA about 19 GB (project table, absolute minimums).
- It covers LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO and FP8 training, and exports to GGUF and other formats (per the project README).
- The speed and memory claims ("2x faster, 70% less VRAM") are the project's own numbers on its own notebooks. Treat them as a ceiling and measure on your model.
- Unsloth Core is Apache-2.0. Unsloth Studio (the UI) and some optional components are AGPL-3.0, which matters if you ship it inside a product.
- Pick it for single-GPU speed and simplicity. Pick Axolotl or LLaMA-Factory for config-driven multi-GPU pipelines, and plain TRL when you need full control of the training loop.
What is Unsloth
Unsloth started as a faster, leaner path through the standard Hugging Face fine-tuning stack. It rewrites the hot parts of training (attention, LoRA math, cross-entropy and others) as optimized kernels so the same fine-tune uses less memory and runs faster. It remains compatible with the Hugging Face ecosystem: the TRL documentation has an official Unsloth integration page, and it describes Unsloth as a framework that trains "up to 2x faster with up to 70% less VRAM".
In 2026 the project is two things:
- Unsloth Core, the Python package. You write or copy a notebook, load a model, attach adapters and train. This is what most tutorials mean by "Unsloth".
- Unsloth Studio, a local web UI and desktop app to run and train models without writing code. Per the README it runs GGUF and MLX models, trains LLMs, diffusion, embedding and audio models, and exposes an OpenAI-compatible API.
The README lists support for current model families, among them Qwen, Gemma 4, DeepSeek-V4, GLM, Kimi and MiniMax models, plus FLUX for image models. It also lists NVIDIA, AMD and Intel GPUs, CPUs and Apple hardware as targets.
What Unsloth can train
From the project README, the training coverage is:
- LoRA and QLoRA (parameter-efficient adapters, see the glossary entries for LoRA and QLoRA)
- Full fine-tuning and pretraining
- Reinforcement learning, including GRPO, with the README singling out GSPO as well
- DPO preference tuning
- FP8 training
Export targets named in the README include GGUF, NVFP4 and FP8. The Unsloth docs say that for Ollama, vLLM or Open WebUI you start from the LoRA adapter plus the base model, that llama.cpp is the route for single-device GGUF inference, and that vLLM is the route for multi-user or FP8 and AWQ deployments. If you are choosing a serving engine afterwards, see serving LLMs with vLLM and the GGUF glossary entry.
On method choice the Unsloth docs recommend starting with QLoRA, and say that with Unsloth's dynamic 4-bit quants the accuracy gap between QLoRA and LoRA is "now largely recovered". That is the project's claim, not an independent measurement.
Install Unsloth
These commands are copied from the Unsloth GitHub README.
Unsloth Studio (macOS, Linux, WSL):
curl -fsSL https://unsloth.ai/install.sh | sh
Windows PowerShell:
irm https://unsloth.ai/install.ps1 | iex
Then launch it:
unsloth studio
The README also documents unsloth studio --secure for public HTTPS access, and a headless password set through the UNSLOTH_STUDIO_PASSWORD environment variable. The install page shows unsloth studio -H 0.0.0.0 -p 8888 for binding the UI to a host and port, which is what you want on a remote machine.
Unsloth Core (code-based, Linux and WSL):
curl -LsSf https://astral.sh/uv/install.sh | sh
uv venv unsloth_env --python 3.13
source unsloth_env/bin/activate
uv pip install unsloth --torch-backend=auto
Docker is covered on the project's Docker install page, with a docker run -d --name unsloth --gpus all --ipc=host -p 8000:8000 -p 8888:8888 ... unsloth/unsloth pattern in the README (the full command, with its volume flags, is on the docs page). AMD users get a separate unsloth/unsloth-rocm image.
Requirements
From the Unsloth requirements page:
- Python 3.11 up to, but not including, 3.14.
- Linux and WSL: Ubuntu 20.04 or similar, with CUDA toolkit 12.4 or newer recommended, and 12.8 or newer for Blackwell GPUs.
- Unsloth Core GPU minimum: CUDA compute capability 7.0 (V100, T4, RTX 20 series and newer, A100, H100, L40 and so on).
- Without a GPU, Unsloth still runs chat and data recipes in Studio.
How much VRAM does Unsloth need
This is the project's own table (source: Unsloth requirements page). QLoRA is 4-bit, LoRA is 16-bit. The page states these are absolute minimums, that some models need more, and suggests batch size 1, 2 or 3 if you run out of memory.
| Model size | QLoRA (4-bit) | LoRA (16-bit) |
|---|---|---|
| 3B | 3.5 GB | 8 GB |
| 7B | 5 GB | 19 GB |
| 8B | 6 GB | 22 GB |
| 14B | 8.5 GB | 33 GB |
| 27B | 22 GB | 64 GB |
| 32B | 26 GB | 76 GB |
| 70B | 41 GB | 164 GB |
| 405B | 237 GB | 950 GB |
Reading it against real cards (computed from the table above): a 24 GB card such as the RTX 4090 fits QLoRA on models up to roughly 14B with room for activations, and LoRA on a 9B model sits at the 24 GB edge. A 48 GB card such as the L40S takes QLoRA on 27B to 32B and 16-bit LoRA on 14B. An 80 GB H100 takes 16-bit LoRA up to about 32B (76 GB, which is tight) or QLoRA on a 70B (41 GB). The project also says a 20B model with more than 500K context can be trained on one 80 GB GPU, per its README news. For your own sizing, use the VRAM calculator for the RTX 4090 or read how much VRAM you need for LLMs. Background on quantization and VRAM is in the glossary.
What speed does Unsloth claim
Everything in this section is published by the Unsloth project, not independently verified by us.
The README headline is training "2x faster with 70% less VRAM" with "no accuracy loss". Its free notebook table lists per-model figures, for example:
| Model | Speed vs baseline | Memory use |
|---|---|---|
| Llama 3.1 (8B) Alpaca | 2x faster | 70% less |
| gpt-oss (20B) | 2x faster | 70% less |
| gpt-oss (20B) GRPO | 2x faster | 80% less |
| Qwen3.5 (4B) | 1.5x faster | 60% less |
| Gemma 4 (E2B) | 1.5x faster | 50% less |
| embeddinggemma (300M) | 2x faster | 20% less |
The README news section adds other claims, such as MoE LLMs trained 12x faster with 35% less VRAM, and 3x faster training with 30% less VRAM in a newer release. These are comparisons against the project's chosen baseline on its notebooks. Your gain depends on model, sequence length and what the baseline already does (Hugging Face now ships its own memory tricks, see the TRL guide linked below). Run one short job with and without Unsloth before you commit.
Unsloth vs Axolotl, LLaMA-Factory and TRL
| Tool | Best at | Interface |
|---|---|---|
| Unsloth | Single-GPU speed and low VRAM, fast start from notebooks | Python notebooks, Studio UI |
| Axolotl | YAML-driven pipelines, multi-GPU and multi-node | YAML plus CLI |
| LLaMA-Factory | Broad model and method coverage, a web UI | YAML plus CLI plus LLaMA Board |
| TRL | Full control, every post-training method, the base others build on | Python trainers |
Unsloth also plugs into TRL, so you can keep TRL's trainers and use Unsloth for the faster model loading and kernels. The direct comparison, including torchtune's status, is in Axolotl vs Unsloth vs torchtune. Single-tool guides: LLaMA-Factory and TRL.
Where Unsloth is a weaker fit: very large multi-node runs, where Axolotl or a DeepSpeed-based stack is the more established path (see DeepSpeed), and any setting where the AGPL terms of Studio are a problem for your distribution model. Check the license of each component you ship.
A practical workflow
- Pick a base model and a method. Start with QLoRA, per the project docs.
- Prepare a few hundred to a few thousand clean examples. Data quality dominates.
- Open a notebook from the Unsloth docs for your model family, or use Studio.
- Train on a small slice first to check the loss falls and the chat template is right.
- Export to GGUF or keep the adapter, then serve it.
The general pipeline, from data to evaluation, is in what AI model fine-tuning is.
Run it on a cloud GPU
Aquanode manages and optimizes GPUs for training and inference workloads. A 24 GB card covers QLoRA up to mid-size models, and an 80 GB card covers 16-bit LoRA on 30B-class models. Live availability and pricing are below.
FAQ
Is Unsloth free?
The Unsloth Core package is Apache-2.0 per the GitHub README, so it is free to use. Studio and some optional components are AGPL-3.0, so read the license before bundling them in a product.
Does Unsloth work on AMD or Apple hardware?
The README lists AMD, Intel, Apple and CPU targets, and ships a separate unsloth/unsloth-rocm Docker image for AMD. Feature parity across hardware is not guaranteed, so check the install page for your platform.
Does Unsloth lower accuracy?
The project claims no accuracy loss for its speedups and says dynamic 4-bit quants largely close the QLoRA versus LoRA gap. Those are the project's claims. Validate on your own evaluation set.
Can I use Unsloth with TRL?
Yes. The TRL documentation has an Unsloth integration page, and Unsloth is designed to work with TRL trainers such as SFTTrainer. See the TRL guide.
What GPU should I start with?
For QLoRA on models up to about 14B, a 24 GB card is enough per the project table. Move to 48 GB for 27B to 32B QLoRA, and 80 GB for 16-bit LoRA at the 30B scale.
Sources
- Unsloth GitHub README: https://github.com/unslothai/unsloth
- Unsloth README news section (12x faster MoE with 35% less VRAM; 500K context for a 20B model on 80 GB; 3x faster with 30% less VRAM): https://raw.githubusercontent.com/unslothai/unsloth/main/README.md
- Unsloth installation docs: https://unsloth.ai/docs/get-started/installing-+-updating
- Unsloth requirements and VRAM table: https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements
- Unsloth fine-tuning guide (LoRA vs QLoRA guidance): https://unsloth.ai/docs/get-started/fine-tuning-llms-guide
- Hugging Face TRL SFT Trainer docs (Unsloth integration note): https://huggingface.co/docs/trl/sft_trainer