TRL Library Guide: SFT, DPO and GRPO Trainers (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

TRL (Transformers Reinforcement Learning) is Hugging Face's library for post-training language models. It gives you trainer classes for supervised fine-tuning (SFTTrainer), preference tuning (DPOTrainer) and reinforcement learning (GRPOTrainer), all built on the Hugging Face Transformers trainer.

This guide uses the TRL documentation, whose pages at the time of writing are for version 1.14.2. Code below is copied from those docs. For the broader tool map, see the LLM fine-tuning frameworks overview, and for the basics, what AI model fine-tuning is.

TL;DR

  • TRL is the code-level layer. A working SFT run is five lines of Python, and the same pattern covers DPO and GRPO.
  • It supports LoRA and QLoRA through PEFT: pass a peft_config and the trainer wraps the model.
  • Unsloth, Axolotl and LLaMA-Factory are higher-level tools, and several of them build on TRL or interoperate with it.
  • GRPOTrainer can use vLLM in colocate mode (inside the trainer process) or server mode (separate GPUs) for faster generation.
  • TRL publishes no per-model VRAM table that we can cite. Size memory with the VRAM calculator and a short test run.

What TRL covers

The docs index organizes the trainers by maturity and type:

  • Online methods: GRPOTrainer, RLOOTrainer
  • Reward modeling: RewardTrainer
  • Offline methods: SFTTrainer, DPOTrainer, KTOTrainer
  • Knowledge distillation: DistillationTrainer
  • Experimental: a longer list including OnlineDPOTrainer, ORPOTrainer, CPOTrainer, GKDTrainer and others

A Hugging Face blog post listed in the docs, "TRL v1: Post-Training Library That Holds When the Field Invalidates Its Own Assumptions" (March 31, 2026), marks the 1.0 line. The docs also include a long-context guide that trains Qwen3-8B on million-token sequences on one 8-GPU node, using sequence parallelism (see the docs page "long context training").

Install is one line, from the installation page:

pip install trl

or with uv:

uv pip install trl

SFTTrainer: supervised fine-tuning

SFT is the simplest and most common way to adapt a model: minimize the negative log-likelihood of the target text given the input. The quickstart from the docs trains Qwen3 0.6B on the Capybara dataset:

from trl import SFTTrainer
from datasets import load_dataset

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()

Dataset formats

SFTTrainer accepts language-modeling and prompt-completion data, each in standard or conversational form. From the docs:

# Standard language modeling
{"text": "The sky is blue."}

# Conversational language modeling
{"messages": [{"role": "user", "content": "What color is the sky?"},
              {"role": "assistant", "content": "It is blue."}]}

# Standard prompt-completion
{"prompt": "The sky is",
 "completion": " blue."}

# Conversational prompt-completion
{"prompt": [{"role": "user", "content": "What color is the sky?"}],
 "completion": [{"role": "assistant", "content": "It is blue."}]}

With a conversational dataset the trainer applies the chat template for you. For prompt-completion data, loss is computed on the completion only by default, and you can set completion_only_loss=False to train on the full sequence. To train only on assistant turns, set assistant_only_loss=True in SFTConfig (the chat template must contain generation markers, which TRL patches automatically for known families such as Qwen3).

Useful SFTConfig options

All of these are documented on the SFT page:

  • packing=True packs several examples into one sequence to cut padding.
  • max_length defaults to 1024 tokens.
  • learning_rate defaults to 2e-5, and the docs suggest roughly 1e-4 when training adapters.
  • gradient_checkpointing defaults to True and bf16 defaults to True unless fp16 is set.
  • use_liger_kernel=True enables Liger Kernel. The docs cite Liger as boosting multi-GPU throughput by 20 percent and cutting memory use by 60 percent, which is Liger's own claim as repeated in the TRL docs.
  • The default loss_type is chunked_nll, which the docs say avoids materializing the full vocabulary-by-sequence logits tensor and so lowers peak activation memory. It is not compatible with Liger, in which case the default becomes nll.

Model precision is set through the model_init_kwargs argument of SFTConfig (for example with a bfloat16 dtype). The docs note that if you do not set a dtype when passing a model id string, it defaults to float32, so set bf16 explicitly to avoid doubling your memory.

LoRA with SFTTrainer

From the docs, adapter training is a single extra argument:

from datasets import load_dataset
from trl import SFTTrainer
from peft import LoraConfig

dataset = load_dataset("trl-lib/Capybara", split="train")

trainer = SFTTrainer(
    "Qwen/Qwen3-0.6B",
    train_dataset=dataset,
    peft_config=LoraConfig(),
)

trainer.train()

For QLoRA, the trainer takes a quantization_config (a BitsAndBytesConfig) alongside peft_config. The docs state to combine the two for QLoRA training. Concepts: LoRA, QLoRA, quantization. For a deeper treatment, see LoRA fine-tuning and PEFT methods.

DPOTrainer: preference tuning

DPO (Direct Preference Optimization, Rafailov et al.) trains on pairs of completions to the same prompt, a preferred one and a rejected one, without a separate reward model. Quickstart from the docs:

from trl import DPOTrainer
from datasets import load_dataset

trainer = DPOTrainer(
    model="Qwen/Qwen3-0.6B",
    train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train"),
)
trainer.train()

The dataset must be a preference dataset. The recommended explicit-prompt shapes are:

# Standard format
preference_example = {"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}

# Conversational format
preference_example = {"prompt": [{"role": "user", "content": "What color is the sky?"}],
                      "chosen": [{"role": "assistant", "content": "It is blue."}],
                      "rejected": [{"role": "assistant", "content": "It is green."}]}

Things to know from the DPOConfig docs:

  • Default learning_rate is 1e-6, much lower than SFT, and the docs suggest about 1e-5 for adapters.
  • Default beta is 0.1. Higher beta means less deviation from the reference model.
  • The default loss is sigmoid, and the docs list many variants (hinge, ipo, bco_pair, sppo_hard, apo_zero, sigmoid_norm and others) selectable through loss_type, including combining several with loss_weights.
  • If you do not pass a ref_model, the trainer uses the initial policy as the reference. That is a second copy of the model in memory unless you use PEFT or precompute_ref_log_probs=True, which the docs say saves memory because the reference model need not stay loaded.

For the theory and when to prefer DPO over PPO, see DPO vs PPO.

GRPOTrainer: reinforcement learning

GRPO (Group Relative Policy Optimization) comes from the DeepSeekMath paper, cited in the TRL docs at https://huggingface.co/papers/2402.03300. You supply reward functions instead of labeled answers, and the trainer samples a group of completions per prompt and learns from relative rewards. The explanation is in GRPO explained. The docs quickstart:

# train_grpo.py
from datasets import load_dataset
from trl import GRPOTrainer
from trl.rewards import accuracy_reward

dataset = load_dataset("trl-lib/DeepMath-103K", split="train")

trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    reward_funcs=accuracy_reward,
    train_dataset=dataset,
)
trainer.train()

Run it with:

accelerate launch train_grpo.py

A custom reward function for a dataset with a ground_truth column, from the docs, returns one float per completion:

import re

def reward_func(completions, ground_truth, **kwargs):
    # Regular expression to capture content inside \boxed{}
    matches = [re.search(r"\\boxed\{(.*?)\}", completion) for completion in completions]
    contents = [match.group(1) if match else "" for match in matches]
    # Reward 1 if the content is the same as the ground truth, 0 otherwise
    return [1.0 if c == gt else 0.0 for c, gt in zip(contents, ground_truth)]

vLLM for generation

RL spends much of its time generating completions, so TRL integrates vLLM. Install with pip install trl[vllm] and set use_vllm=True in GRPOConfig. Two modes, per the docs:

  • Colocate (the default, vllm_mode="colocate"): vLLM runs inside the trainer process and shares the GPUs with training. vllm_gpu_memory_utilization defaults to 0.3, and vllm_enable_sleep_mode=True offloads vLLM weights and cache during the optimizer step.
  • Server (vllm_mode="server"): vLLM runs as a separate process on dedicated GPUs. The docs warn to use different GPUs for the server and the trainer (via CUDA_VISIBLE_DEVICES) to avoid NCCL errors.

Server start command from the docs:

VLLM_SERVER_DEV_MODE=1 vllm serve <model_name> \
    --weight-transfer-config '{"backend": "nccl"}' \
    --logprobs-mode processed_logprobs \
    --max-logprobs -1

Truncated importance sampling is on by default to correct the mismatch between the generation engine and the training model. For vLLM itself, see serving LLMs with vLLM.

Scaling TRL beyond one GPU

TRL trainers inherit the Transformers trainer, so multi-GPU runs go through Accelerate, DeepSpeed or FSDP. SFTConfig and DPOConfig expose fsdp, fsdp_config and deepspeed arguments. Background: FSDP and the DeepSpeed guide. The docs integrate DeepSpeed, Liger Kernel and PEFT as first-class options.

TRL vs Unsloth, Axolotl and LLaMA-Factory

ToolLevelBest for
TRLPython libraryCustom loops, research, every post-training method, the reference implementation
UnslothOptimized training layerSingle-GPU speed and low VRAM; the TRL docs themselves describe a documented integration
AxolotlYAML pipelineRepeatable multi-GPU runs
LLaMA-FactoryCLI and web UIMany models and methods, no code

Choose TRL when you need to change the loss, write your own reward, mix methods, or embed training inside a larger Python system. Choose a wrapper when the defaults fit and you want less glue. Compare: Unsloth guide, LLaMA-Factory guide, Axolotl vs Unsloth vs torchtune.

Sizing the GPU

TRL's docs give no per-model memory table, so here is a computed rule of thumb, not a TRL figure: bf16 weights take about 2 bytes per parameter, so a 7B model is about 14 GB before gradients, optimizer state and activations. LoRA avoids most of the gradient and optimizer cost, which is why a 7B LoRA run fits on a 24 GB card while full fine-tuning does not. GRPO adds generation memory (KV cache), and DPO adds a reference model unless you use PEFT or precomputed reference log probabilities. See how much VRAM you need for LLMs and the KV cache entry.

Run it on a cloud GPU

Aquanode manages and optimizes GPUs for training and inference workloads. SFT and DPO with LoRA on small and mid-size models fit a 48 GB card, while GRPO with colocated vLLM and larger models is more comfortable on 80 GB. Live availability and pricing are below.

FAQ

What does TRL stand for?

Transformers Reinforcement Learning. The docs describe it as a full-stack library for training transformer language models with SFT, GRPO, DPO, reward modeling and more.

Does TRL support LoRA and QLoRA?

Yes, through PEFT. Pass peft_config=LoraConfig() to the trainer, and add a quantization_config for QLoRA.

SFTTrainer or DPOTrainer first?

Start with SFT to teach the format and domain, then use DPO if you have preference pairs and want to shift style or behavior. DPO usually runs at a much lower learning rate.

Can TRL train vision-language models?

The SFT and DPO docs state both trainers support vision-language models, using a dataset with an image or images column, and recommend setting max_length=None so image tokens are not truncated.

How is GRPO different from DPO?

DPO learns from fixed chosen and rejected pairs. GRPO samples completions during training and learns from reward functions you define, so it needs generation and benefits from vLLM.

Sources

#fine-tuning#llm training#trl#sft#dpo#grpo

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.