TRL (Transformers Reinforcement Learning) is Hugging Face's library for post-training language models. It gives you trainer classes for supervised fine-tuning (SFTTrainer), preference tuning (DPOTrainer) and reinforcement learning (GRPOTrainer), all built on the Hugging Face Transformers trainer.
This guide uses the TRL documentation, whose pages at the time of writing are for version 1.14.2. Code below is copied from those docs. For the broader tool map, see the LLM fine-tuning frameworks overview, and for the basics, what AI model fine-tuning is.
TL;DR
- TRL is the code-level layer. A working SFT run is five lines of Python, and the same pattern covers DPO and GRPO.
- It supports LoRA and QLoRA through PEFT: pass a
peft_configand the trainer wraps the model. - Unsloth, Axolotl and LLaMA-Factory are higher-level tools, and several of them build on TRL or interoperate with it.
- GRPOTrainer can use vLLM in colocate mode (inside the trainer process) or server mode (separate GPUs) for faster generation.
- TRL publishes no per-model VRAM table that we can cite. Size memory with the VRAM calculator and a short test run.
What TRL covers
The docs index organizes the trainers by maturity and type:
- Online methods: GRPOTrainer, RLOOTrainer
- Reward modeling: RewardTrainer
- Offline methods: SFTTrainer, DPOTrainer, KTOTrainer
- Knowledge distillation: DistillationTrainer
- Experimental: a longer list including OnlineDPOTrainer, ORPOTrainer, CPOTrainer, GKDTrainer and others
A Hugging Face blog post listed in the docs, "TRL v1: Post-Training Library That Holds When the Field Invalidates Its Own Assumptions" (March 31, 2026), marks the 1.0 line. The docs also include a long-context guide that trains Qwen3-8B on million-token sequences on one 8-GPU node, using sequence parallelism (see the docs page "long context training").
Install is one line, from the installation page:
pip install trl
or with uv:
uv pip install trl
SFTTrainer: supervised fine-tuning
SFT is the simplest and most common way to adapt a model: minimize the negative log-likelihood of the target text given the input. The quickstart from the docs trains Qwen3 0.6B on the Capybara dataset:
from trl import SFTTrainer
from datasets import load_dataset
trainer = SFTTrainer(
model="Qwen/Qwen3-0.6B",
train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()
Dataset formats
SFTTrainer accepts language-modeling and prompt-completion data, each in standard or conversational form. From the docs:
# Standard language modeling
{"text": "The sky is blue."}
# Conversational language modeling
{"messages": [{"role": "user", "content": "What color is the sky?"},
{"role": "assistant", "content": "It is blue."}]}
# Standard prompt-completion
{"prompt": "The sky is",
"completion": " blue."}
# Conversational prompt-completion
{"prompt": [{"role": "user", "content": "What color is the sky?"}],
"completion": [{"role": "assistant", "content": "It is blue."}]}
With a conversational dataset the trainer applies the chat template for you. For prompt-completion data, loss is computed on the completion only by default, and you can set completion_only_loss=False to train on the full sequence. To train only on assistant turns, set assistant_only_loss=True in SFTConfig (the chat template must contain generation markers, which TRL patches automatically for known families such as Qwen3).
Useful SFTConfig options
All of these are documented on the SFT page:
packing=Truepacks several examples into one sequence to cut padding.max_lengthdefaults to 1024 tokens.learning_ratedefaults to 2e-5, and the docs suggest roughly 1e-4 when training adapters.gradient_checkpointingdefaults to True andbf16defaults to True unlessfp16is set.use_liger_kernel=Trueenables Liger Kernel. The docs cite Liger as boosting multi-GPU throughput by 20 percent and cutting memory use by 60 percent, which is Liger's own claim as repeated in the TRL docs.- The default
loss_typeischunked_nll, which the docs say avoids materializing the full vocabulary-by-sequence logits tensor and so lowers peak activation memory. It is not compatible with Liger, in which case the default becomesnll.
Model precision is set through the model_init_kwargs argument of SFTConfig (for example with a bfloat16 dtype). The docs note that if you do not set a dtype when passing a model id string, it defaults to float32, so set bf16 explicitly to avoid doubling your memory.
LoRA with SFTTrainer
From the docs, adapter training is a single extra argument:
from datasets import load_dataset
from trl import SFTTrainer
from peft import LoraConfig
dataset = load_dataset("trl-lib/Capybara", split="train")
trainer = SFTTrainer(
"Qwen/Qwen3-0.6B",
train_dataset=dataset,
peft_config=LoraConfig(),
)
trainer.train()
For QLoRA, the trainer takes a quantization_config (a BitsAndBytesConfig) alongside peft_config. The docs state to combine the two for QLoRA training. Concepts: LoRA, QLoRA, quantization. For a deeper treatment, see LoRA fine-tuning and PEFT methods.
DPOTrainer: preference tuning
DPO (Direct Preference Optimization, Rafailov et al.) trains on pairs of completions to the same prompt, a preferred one and a rejected one, without a separate reward model. Quickstart from the docs:
from trl import DPOTrainer
from datasets import load_dataset
trainer = DPOTrainer(
model="Qwen/Qwen3-0.6B",
train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train"),
)
trainer.train()
The dataset must be a preference dataset. The recommended explicit-prompt shapes are:
# Standard format
preference_example = {"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}
# Conversational format
preference_example = {"prompt": [{"role": "user", "content": "What color is the sky?"}],
"chosen": [{"role": "assistant", "content": "It is blue."}],
"rejected": [{"role": "assistant", "content": "It is green."}]}
Things to know from the DPOConfig docs:
- Default
learning_rateis 1e-6, much lower than SFT, and the docs suggest about 1e-5 for adapters. - Default
betais 0.1. Higher beta means less deviation from the reference model. - The default loss is
sigmoid, and the docs list many variants (hinge, ipo, bco_pair, sppo_hard, apo_zero, sigmoid_norm and others) selectable throughloss_type, including combining several withloss_weights. - If you do not pass a
ref_model, the trainer uses the initial policy as the reference. That is a second copy of the model in memory unless you use PEFT orprecompute_ref_log_probs=True, which the docs say saves memory because the reference model need not stay loaded.
For the theory and when to prefer DPO over PPO, see DPO vs PPO.
GRPOTrainer: reinforcement learning
GRPO (Group Relative Policy Optimization) comes from the DeepSeekMath paper, cited in the TRL docs at https://huggingface.co/papers/2402.03300. You supply reward functions instead of labeled answers, and the trainer samples a group of completions per prompt and learns from relative rewards. The explanation is in GRPO explained. The docs quickstart:
# train_grpo.py
from datasets import load_dataset
from trl import GRPOTrainer
from trl.rewards import accuracy_reward
dataset = load_dataset("trl-lib/DeepMath-103K", split="train")
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-0.5B-Instruct",
reward_funcs=accuracy_reward,
train_dataset=dataset,
)
trainer.train()
Run it with:
accelerate launch train_grpo.py
A custom reward function for a dataset with a ground_truth column, from the docs, returns one float per completion:
import re
def reward_func(completions, ground_truth, **kwargs):
# Regular expression to capture content inside \boxed{}
matches = [re.search(r"\\boxed\{(.*?)\}", completion) for completion in completions]
contents = [match.group(1) if match else "" for match in matches]
# Reward 1 if the content is the same as the ground truth, 0 otherwise
return [1.0 if c == gt else 0.0 for c, gt in zip(contents, ground_truth)]
vLLM for generation
RL spends much of its time generating completions, so TRL integrates vLLM. Install with pip install trl[vllm] and set use_vllm=True in GRPOConfig. Two modes, per the docs:
- Colocate (the default,
vllm_mode="colocate"): vLLM runs inside the trainer process and shares the GPUs with training.vllm_gpu_memory_utilizationdefaults to 0.3, andvllm_enable_sleep_mode=Trueoffloads vLLM weights and cache during the optimizer step. - Server (
vllm_mode="server"): vLLM runs as a separate process on dedicated GPUs. The docs warn to use different GPUs for the server and the trainer (viaCUDA_VISIBLE_DEVICES) to avoid NCCL errors.
Server start command from the docs:
VLLM_SERVER_DEV_MODE=1 vllm serve <model_name> \
--weight-transfer-config '{"backend": "nccl"}' \
--logprobs-mode processed_logprobs \
--max-logprobs -1
Truncated importance sampling is on by default to correct the mismatch between the generation engine and the training model. For vLLM itself, see serving LLMs with vLLM.
Scaling TRL beyond one GPU
TRL trainers inherit the Transformers trainer, so multi-GPU runs go through Accelerate, DeepSpeed or FSDP. SFTConfig and DPOConfig expose fsdp, fsdp_config and deepspeed arguments. Background: FSDP and the DeepSpeed guide. The docs integrate DeepSpeed, Liger Kernel and PEFT as first-class options.
TRL vs Unsloth, Axolotl and LLaMA-Factory
| Tool | Level | Best for |
|---|---|---|
| TRL | Python library | Custom loops, research, every post-training method, the reference implementation |
| Unsloth | Optimized training layer | Single-GPU speed and low VRAM; the TRL docs themselves describe a documented integration |
| Axolotl | YAML pipeline | Repeatable multi-GPU runs |
| LLaMA-Factory | CLI and web UI | Many models and methods, no code |
Choose TRL when you need to change the loss, write your own reward, mix methods, or embed training inside a larger Python system. Choose a wrapper when the defaults fit and you want less glue. Compare: Unsloth guide, LLaMA-Factory guide, Axolotl vs Unsloth vs torchtune.
Sizing the GPU
TRL's docs give no per-model memory table, so here is a computed rule of thumb, not a TRL figure: bf16 weights take about 2 bytes per parameter, so a 7B model is about 14 GB before gradients, optimizer state and activations. LoRA avoids most of the gradient and optimizer cost, which is why a 7B LoRA run fits on a 24 GB card while full fine-tuning does not. GRPO adds generation memory (KV cache), and DPO adds a reference model unless you use PEFT or precomputed reference log probabilities. See how much VRAM you need for LLMs and the KV cache entry.
Run it on a cloud GPU
Aquanode manages and optimizes GPUs for training and inference workloads. SFT and DPO with LoRA on small and mid-size models fit a 48 GB card, while GRPO with colocated vLLM and larger models is more comfortable on 80 GB. Live availability and pricing are below.
FAQ
What does TRL stand for?
Transformers Reinforcement Learning. The docs describe it as a full-stack library for training transformer language models with SFT, GRPO, DPO, reward modeling and more.
Does TRL support LoRA and QLoRA?
Yes, through PEFT. Pass peft_config=LoraConfig() to the trainer, and add a quantization_config for QLoRA.
SFTTrainer or DPOTrainer first?
Start with SFT to teach the format and domain, then use DPO if you have preference pairs and want to shift style or behavior. DPO usually runs at a much lower learning rate.
Can TRL train vision-language models?
The SFT and DPO docs state both trainers support vision-language models, using a dataset with an image or images column, and recommend setting max_length=None so image tokens are not truncated.
How is GRPO different from DPO?
DPO learns from fixed chosen and rejected pairs. GRPO samples completions during training and learns from reward functions you define, so it needs generation and benefits from vLLM.
Sources
- TRL documentation index: https://huggingface.co/docs/trl/index
- TRL installation: https://huggingface.co/docs/trl/installation
- SFT Trainer docs: https://huggingface.co/docs/trl/sft_trainer
- DPO Trainer docs: https://huggingface.co/docs/trl/dpo_trainer
- GRPO Trainer docs: https://huggingface.co/docs/trl/grpo_trainer
- DPO paper (Rafailov et al.): https://huggingface.co/papers/2305.18290
- DeepSeekMath paper (GRPO): https://huggingface.co/papers/2402.03300