PPO-based RLHF trains a separate reward model from human preferences and then runs reinforcement learning against it, sampling from the model as it trains. DPO skips both steps: it optimizes the policy directly on pairs of preferred and rejected answers with a simple classification-style loss. The DPO paper reports that it matches or beats PPO-based RLHF on the tasks it tested, with far less machinery.
This guide explains both, shows what each costs in GPU memory with the math, gives the TRL code for DPO, and tells you when PPO-style online training (including GRPO) is still the right call. It belongs to our LLM fine-tuning frameworks guide.
TL;DR
- RLHF (reinforcement learning from human feedback) is the three-stage recipe: supervised fine-tuning, reward model training on human comparisons, then RL (PPO) against that reward model.
- PPO is the RL step. It samples from the model during training and needs a policy, a reference model, a reward model and a value model.
- DPO reparameterizes the reward so the optimal policy has a closed form, then trains on preference pairs with one loss. No reward model, no sampling, no value model.
- Choose DPO when you have a preference dataset and want the simplest stable pipeline. Choose online RL (PPO, or GRPO) when you can score fresh model outputs, for example with a verifier.
- DPO still needs the reference model's log-probabilities, so it holds roughly two models in memory, not one.
What RLHF is
The InstructGPT paper (Ouyang et al., 2022) set the template. In its words, the pipeline is:
- Collect demonstration data and train a supervised policy (fine-tune a pre-trained model with supervised learning).
- Collect comparison data, where labelers indicate which output they prefer, and train a reward model to predict the preferred output.
- Optimize a policy against the reward model using PPO, using the reward model output as a scalar reward.
Steps 2 and 3 can be iterated. A per-token KL penalty from the supervised model is added at each token to mitigate over-optimization of the reward model. The headline result: in human evaluations on the paper's prompt distribution, outputs from the 1.3B InstructGPT model were preferred to outputs from the 175B GPT-3, despite having 100 times fewer parameters. The paper's takeaway is that making models bigger does not inherently make them better at following a user's intent.
Note a detail that matters for memory below: InstructGPT used a single 6B reward model, and a 6B value function, for all PPO model sizes, saving compute. Nothing forces the critic to match the policy.
For why post-training exists at all, see what AI model fine-tuning involves.
How PPO works for LLMs
PPO (Schulman et al., 2017) is a family of policy gradient methods that alternates between collecting data by interacting with the environment and optimizing a surrogate objective with several epochs of minibatch updates, instead of one update per sample. In the LLM setting, the "environment" presents a prompt, the model produces a response, and the reward model scores it.
The moving parts, per the InstructGPT and DeepSeekMath descriptions:
- Policy: the model being trained.
- Reference model: the frozen starting point used for the KL penalty.
- Reward model: scores completions.
- Value model (critic): estimates the baseline for advantage computation. InstructGPT initialized its value function from the reward model.
The DeepSeekMath authors point out that the value function is typically another model of comparable size and brings a substantial memory and computational burden, and that because usually only the last token gets a reward, a token-accurate value function is hard to train. That observation is what led to GRPO, which removes the critic.
PPO is powerful because it is online: the policy generates its own outputs and learns from the reward on them. The price is complexity and four models.
How DPO works
The DPO paper (Rafailov et al., 2023) opens with the problem: RLHF is a complex and often unstable procedure, first fitting a reward model and then fine-tuning with RL to maximize it without drifting too far from the original model. DPO introduces a new parameterization of the reward model that allows the optimal policy to be extracted in closed form, so the standard RLHF problem can be solved with only a simple classification loss. The authors describe the result as stable, performant and computationally lightweight, eliminating the need for sampling from the LM during fine-tuning or significant hyperparameter tuning.
The loss, as documented by TRL (written without math markup):
loss = - log sigmoid( beta * ( log(pi(chosen) / pi_ref(chosen))
- log(pi(rejected) / pi_ref(rejected)) ) )
Here pi is the policy being trained, pi_ref is the reference model, and beta controls the strength of the preference signal. TRL describes the effect as widening the margin between the log-likelihoods of preferred and dispreferred completions, relative to the reference model, and notes that in practice this is usually achieved by suppressing the dispreferred completions rather than raising the preferred ones.
The quantity beta * log(pi / pi_ref) is the implicit reward. TRL logs it as rewards/chosen and rewards/rejected, plus rewards/margins and rewards/accuracies, which is how you monitor a DPO run without a separate reward model.
What the paper claims
From the abstract: DPO can fine-tune language models to align with human preferences as well as or better than existing methods. Notably, fine-tuning with DPO exceeds PPO-based RLHF in the ability to control the sentiment of generations, and matches or improves response quality in summarization and single-turn dialogue, while being substantially simpler to implement and train. Those results are from the paper's own experiments and model sizes. They are not a guarantee for your task, and later work has proposed variants (below) because plain DPO has known quirks such as length bias that TRL addresses with a length-normalized loss option.
DPO vs PPO side by side
| PPO (RLHF) | DPO | |
|---|---|---|
| Training data | Prompts, plus a reward model trained on comparisons | Preference pairs: prompt, chosen, rejected |
| Samples from the model while training | Yes | No |
| Reward model | Required | Not needed (implicit reward) |
| Value model | Required | Not needed |
| Reference model | Required (KL penalty) | Required (log-ratios) |
| Moving parts | Four models, RL loop | One loss, two models |
| Can learn from fresh outputs | Yes | Only what is in the dataset |
| Typical use | Online alignment, reward-driven training | Preference alignment from a fixed dataset |
Memory cost (computed)
Using 16 bytes per trainable parameter for full fine-tuning model states (2 weights, 2 gradients, 12 mixed-precision Adam, from the ZeRO paper) and 2 bytes per parameter for a frozen BF16 model. Activations are extra.
For a 7B model, full fine-tuning:
PPO: policy 16*7B = 112 GB + value model 112 GB (if policy-sized)
+ reference 2*7B = 14 GB + reward model 14 GB = 252 GB
DPO: policy 112 GB + reference 14 GB = 126 GB
DPO with precomputed reference log-probs = 112 GB
Assumptions: the PPO value model is policy-sized and trained with the same recipe, and the reward model is policy-sized. InstructGPT used smaller 6B critics and reward models for every policy size, so a real PPO stack can be lighter. The PPO figure here is an upper bound for a symmetric setup. TRL's precompute_ref_log_probs option computes the reference log-probabilities for the whole dataset before training, so the reference model does not need to stay in memory.
On GPUs, 126 GB of states exceeds one 80 GB H100 and sits inside one 141 GB H200 only before activations, so full DPO at 7B is a multi-GPU job in practice (see FSDP). With adapters the picture changes: freeze the base in BF16 or 4-bit and train a LoRA adapter, as in the LoRA fine-tuning guide. A 7B BF16 base is 14 GB and a 4-bit base is 3.5 GB (computed), both single-GPU territory.
DPO with TRL
The quickstart from the TRL docs (main branch):
from trl import DPOTrainer
from datasets import load_dataset
trainer = DPOTrainer(
model="Qwen/Qwen3-0.6B",
train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train"),
)
trainer.train()
The dataset must be a preference dataset. TRL accepts an explicit prompt (recommended) or implicit prompt, in standard or conversational form:
preference_example = {"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}
When you do not pass a reference model, TRL uses the initial policy, the model before DPO training starts. To train adapters instead of all weights, pass a PEFT config:
from datasets import load_dataset
from trl import DPOTrainer
from peft import LoraConfig
dataset = load_dataset("trl-lib/ultrafeedback_binarized", split="train")
trainer = DPOTrainer(
"Qwen/Qwen3-0.6B",
train_dataset=dataset,
peft_config=LoraConfig(),
)
trainer.train()
Defaults worth knowing from DPOConfig: beta=0.1 (a higher beta means less deviation from the reference), learning_rate=1e-6, max_length=1024, loss_type=["sigmoid"]. TRL's tip for adapters is a higher learning rate, around 1e-5. A common recipe is to run supervised fine-tuning first, then DPO from that checkpoint; see the TRL library guide for the trainer family.
Loss variants in TRL
TRL's loss_type exposes variants proposed after the original paper, each tied to a paper in the docs. A few: "hinge" (RSO), "ipo" (IPO, which argues the logit transform can overfit), "robust" (unbiased under noisy preferences), "bco_pair", "sigmoid_norm" (SimPO, normalizing by token count to address length bias in the original loss), and "sft". You can combine them with loss_weights, as in TRL's MPO example. Start with the default and change one thing at a time.
When to use which
Pick DPO when:
- You have, or can build, a preference dataset (chosen versus rejected pairs).
- You want a stable, single-stage run on modest hardware.
- The goal is style, helpfulness, tone or safety behavior that humans judge.
Pick online RL (PPO or GRPO) when:
- You can score fresh outputs automatically, for example math answers or code with tests. DeepSeek-R1-Zero used rule-based accuracy and format rewards with GRPO and no human-labeled reasoning traces. See GRPO explained.
- You need the model to improve on its own generations rather than a fixed dataset.
- You already have a trustworthy reward model.
Combine them: nothing prevents SFT, then DPO for preferences, then a verifiable-reward RL stage. Each is a separate TRL trainer, and each accepts adapters, so the PEFT methods guide applies to all.
Run it on a cloud GPU
DPO and RL runs keep a reference model and long sequences in memory, so favor large-memory cards. The live box shows what is rentable right now.
FAQ
Is DPO better than PPO?
Not universally. The DPO paper reports it matches or exceeds PPO-based RLHF on sentiment control, summarization and single-turn dialogue, with simpler training. PPO and its variants remain the choice when you need online learning from a reward signal.
Does DPO need a reward model?
No. The reward is implicit, beta * log(pi / pi_ref). You need preference pairs and a reference model.
What is RLHF in simple terms?
Train a model on human demonstrations, train a reward model on human comparisons of outputs, then use reinforcement learning to push the model toward higher reward while a KL penalty keeps it close to where it started.
What dataset format does DPO need?
Prompt, chosen and rejected fields, in standard or conversational form. TRL's trl-lib/ultrafeedback_binarized is the dataset in its quickstart.
How much VRAM does DPO need?
Computed for a 7B model: about 126 GB of model states in full fine-tuning (112 GB policy plus 14 GB BF16 reference), and about 112 GB if you precompute reference log-probabilities. With a LoRA adapter over a BF16 base, the base is 14 GB plus activations.
Is GRPO a replacement for PPO?
It is a PPO variant that removes the value model and uses the group average reward as the baseline. DeepSeek presented it as optimizing PPO's memory usage. It still needs a reward signal and sampling.
Sources
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al.): https://arxiv.org/abs/2305.18290
- Training language models to follow instructions with human feedback (Ouyang et al., InstructGPT): https://arxiv.org/abs/2203.02155
- Proximal Policy Optimization Algorithms (Schulman et al.): https://arxiv.org/abs/1707.06347
- DeepSeekMath (value model burden, GRPO): https://arxiv.org/abs/2402.03300
- DeepSeek-R1 (rule-based rewards, GRPO): https://arxiv.org/abs/2501.12948
- ZeRO (16 bytes per parameter breakdown): https://arxiv.org/abs/1910.02054
- TRL DPOTrainer docs: https://huggingface.co/docs/trl/main/en/dpo_trainer
- TRL GRPOTrainer docs: https://huggingface.co/docs/trl/main/en/grpo_trainer
- NVIDIA H100: https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200: https://www.nvidia.com/en-us/data-center/h200/