GRPO (Group Relative Policy Optimization) is a reinforcement learning algorithm for language models that replaces PPO's learned value network with a simple statistic: it samples a group of answers to the same prompt, scores them, and uses each answer's reward relative to the group as its advantage. DeepSeek introduced it in DeepSeekMath and used it to train DeepSeek-R1-Zero.
This guide explains the mechanism with worked numbers, lists the hyperparameters the papers actually used, shows the Hugging Face TRL code, and computes what it costs in GPU memory. It is part of our LLM fine-tuning frameworks guide.
TL;DR
- GRPO is a variant of PPO. It drops the critic (value model) and estimates the baseline from the average reward of a group of sampled outputs for the same prompt.
- The advantage of each output is its reward minus the group mean, divided by the group standard deviation.
- It works best when you can check answers automatically (math, code, formats), because rewards can be plain functions rather than a trained reward model.
- Memory is lower than PPO because there is no second model of policy size to train. TRL's default also skips the reference model.
- If every sample in a group gets the same reward, that group teaches nothing. Reward design matters more than any other knob.
For how it compares with preference methods, read DPO vs PPO. To train the policy with adapters, see the LoRA fine-tuning guide.
Where GRPO comes from
PPO, introduced by Schulman et al. in 2017, is the reinforcement learning algorithm used in the RL stage of many LLM training pipelines (the InstructGPT paper is the best-known example). In LLM use, PPO needs a policy, a reference model for the KL penalty, a reward model, and a value model.
The DeepSeekMath paper (Shao et al., 2024) describes the problem with the value model: it is typically another model of comparable size to the policy, which brings a substantial memory and computational burden. They add that in the LLM setting usually only the last token receives a reward score, which complicates training a value function that is accurate at each token. GRPO obviates the value function and instead uses the average reward of multiple sampled outputs, produced in response to the same question, as the baseline.
The abstract presents GRPO as a variant of PPO that enhances mathematical reasoning while optimizing PPO's memory usage. In that paper, DeepSeekMath 7B reaches 51.7% on the competition-level MATH benchmark without external toolkits or voting, and 60.9% with self-consistency over 64 samples. Those are the authors' reported results.
The mechanism, step by step
For each training prompt q:
- Sample a group. Draw G outputs from the current (old) policy. The paper calls the group
o_1 ... o_G. - Score each output. A reward model, or in the R1 setup a rule-based check, produces rewards
r_1 ... r_G. - Normalize inside the group. Subtract the group average and divide by the group standard deviation. For outcome supervision, every token in output i gets that normalized reward as its advantage.
- Update the policy. Maximize a PPO-style clipped objective weighted by those advantages, minus a KL penalty to a reference policy.
The advantage, in the paper's notation written out in plain text:
A_i = (r_i - mean(r_1..r_G)) / std(r_1..r_G)
And the objective for one output, per the paper (per-token ratio against the old policy, clipped to 1 minus epsilon and 1 plus epsilon):
J = mean over outputs and tokens of
min( ratio * A, clip(ratio, 1 - eps, 1 + eps) * A )
- beta * KL(policy || reference)
A worked example (computed)
Suppose G = 4 and a verifier returns 1 for a correct answer and 0 for a wrong one. These numbers use the population standard deviation; libraries may use the sample version, which changes the scale slightly but not the signs.
| Group rewards | Mean | Std | Advantages |
|---|---|---|---|
| 1, 0, 0, 1 | 0.5 | 0.5 | +1, -1, -1, +1 |
| 1, 0, 0, 0 | 0.25 | 0.433 | +1.73, -0.58, -0.58, -0.58 |
| 0, 0, 0, 0 | 0 | 0 | all 0 (numerator is zero) |
Read the second row: one lucky correct answer out of four gets a large positive push, and the three wrong ones get a mild negative push. In the third row every output scored the same, so every advantage is zero and the group contributes no gradient. This is why prompts that are always solved or never solved are wasted compute, and why the next section's hyperparameters matter.
What the papers actually used
These are reported settings, not recommendations.
| Setting | DeepSeekMath RL (7B) | DeepSeek-R1-Zero |
|---|---|---|
| Policy learning rate | 1e-6 | 3e-6 |
| KL coefficient | 0.04 | 0.001 |
| Outputs sampled per question | 64 | 16 |
| Max output length | 1024 | 32,768 tokens, then 65,536 after step 8.2k |
| Batch | 1024 | 32 unique questions, 512 outputs per step |
DeepSeek-R1-Zero rewards, per the R1 paper, are rule-based and come in two types: accuracy rewards (is the response correct, checked by rules) and format rewards (enforcing a specified reasoning format), combined with the same weight. The abstract of the R1 paper states the claim that reasoning abilities can be incentivized through pure reinforcement learning, obviating human-labeled reasoning trajectories. The paper's text reports the average AIME 2024 pass@1 of R1-Zero rising from 15.6% to 77.9% during RL training (this figure is from the January 2026 revision, arXiv v2). The authors also describe a rollout of 8,192 outputs per step, split into 16 minibatches with a single inner epoch, and replacing the reference model with the latest policy every 400 steps.
GRPO versus PPO versus DPO
| PPO | GRPO | DPO | |
|---|---|---|---|
| Needs a value model | Yes | No | No |
| Needs a reward model or verifier | Yes | Yes (can be a rule) | No, uses preference pairs |
| Samples from the model during training | Yes | Yes (a group per prompt) | No |
| Data | Prompts plus reward signal | Prompts plus reward signal | Chosen and rejected pairs |
GRPO sits in the middle: online like PPO, simpler than PPO, and driven by verifiable rewards instead of human preference labels. More on the other two in DPO vs PPO.
Memory: what you need on the GPU
All values below are computed. The formula for full fine-tuning model states is 16 bytes per trainable parameter (2 weights, 2 gradients, 12 for mixed-precision Adam, from the ZeRO paper), and 2 bytes per parameter for a frozen BF16 model. Activations and generation KV cache are extra.
For a 7B policy:
PPO: policy 16*7B = 112 GB + value (same size) 112 GB
+ reference 2*7B = 14 GB + reward model 14 GB = 252 GB
GRPO: policy 112 GB + reference 14 GB (only if beta > 0) = 126 GB
policy only when beta = 0 = 112 GB
The PPO line assumes the value model is policy-sized and trained with the same recipe, which matches the DeepSeekMath description, and a reward model of the same size. Real setups vary. The point is the direction: GRPO removes the second trained network, roughly halving the model-state memory.
Generation is the other cost. GRPO samples G outputs per prompt and long reasoning traces make the KV cache large, so TRL integrates vLLM for rollouts. Even 112 GB of full-precision states exceeds one 80 GB card, so for single-GPU work use LoRA or QLoRA on the policy (see QLoRA and quantization). Thinking Machines' "LoRA Without Regret" reports that LoRA fully matches full fine-tuning in policy-gradient RL even at rank 1, attributing it to RL absorbing little information per episode. Sizing guidance for cards is in the VRAM sizing guide.
Training GRPO with TRL
TRL's GRPOTrainer is the reference implementation most teams use. This quickstart is from the TRL docs (main branch):
# train_grpo.py
from datasets import load_dataset
from trl import GRPOTrainer
from trl.rewards import accuracy_reward
dataset = load_dataset("trl-lib/DeepMath-103K", split="train")
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-0.5B-Instruct",
reward_funcs=accuracy_reward,
train_dataset=dataset,
)
trainer.train()
accelerate launch train_grpo.py
Key GRPOConfig defaults from the same page:
| Parameter | Default | Meaning |
|---|---|---|
num_generations | 8 | Completions sampled per prompt (G). The effective batch size must be divisible by it. |
beta | 0.0 | KL coefficient. At 0.0 the reference model is not loaded. |
loss_type | "dapo" | Other documented options: "dr_grpo" and "sapo". |
learning_rate | 1e-6 | Same order as the DeepSeekMath setting. |
Rewards can be a built-in function, your own Python function, a model ID, or a list. A custom function receives the prompts and completions plus dataset columns and returns one float per completion. With several reward functions the values are summed, or weighted with reward_weights. TRL normalizes rewards inside each prompt's group to get the advantages, matching the paper. The docs' example reward:
def reward_func(completion_ids, **kwargs):
"""Reward longer completions (in tokens)."""
return [float(len(ids)) for ids in completion_ids]
That one is a toy. Real reward functions check an answer, run unit tests, or validate a format, as in the R1 setup.
For fast rollouts, install the extra with pip install trl[vllm] and set use_vllm=True. Colocate mode (the default) runs vLLM inside the trainer process and shares the GPUs. Server mode (vllm_mode="server") runs vllm serve MODEL on separate GPUs. Background on the engine is in serving LLMs with vLLM.
If you prefer a higher-level wrapper, the TRL library guide covers the trainer family and the Unsloth guide covers a memory-optimized path.
Practical pitfalls
- Dead groups. From the table above, identical rewards give zero advantage. Filter prompts that are always solved or never solved, or choose a harder or easier dataset.
- Reward hacking. A model optimizes whatever the reward function measures. A format reward alone can be satisfied without any reasoning, so check samples by hand and make sure the accuracy term carries the signal.
- KL setting. The two papers differ by a factor of 40 in KL coefficient (0.04 versus 0.001), and TRL defaults to zero. Treat it as something to tune and monitor.
- Length. R1-Zero raised the max length from 32,768 to 65,536 tokens mid-run, which the paper links to a jump in performance and response length. Rollout length dominates your memory and time budget.
Run it on a cloud GPU
GRPO is generation-heavy, so favor large-memory cards. The live box shows what is rentable right now.
FAQ
What does GRPO stand for?
Group Relative Policy Optimization. "Group" refers to the set of outputs sampled for each prompt, and "relative" to scoring each output against the group average.
How is GRPO different from PPO?
GRPO has no value model. PPO learns a critic to estimate the baseline; GRPO uses the mean reward of the sampled group instead, which saves a policy-sized network.
Do I need a reward model for GRPO?
Not necessarily. DeepSeek-R1-Zero used rule-based accuracy and format rewards. In TRL a reward can be any Python function that returns a float per completion.
How many generations per prompt should I use?
The papers used 64 (DeepSeekMath) and 16 (R1-Zero). TRL's default is 8. More samples give a better group baseline but multiply generation cost.
Can I run GRPO with LoRA?
Yes. TRL trainers accept a PEFT config, and the "LoRA Without Regret" study reports LoRA matching full fine-tuning in RL even at low rank. See the PEFT methods guide.
Is GRPO only for math?
No, but it fits tasks with checkable answers: math, code with tests, structured outputs. For subjective quality, preference methods like DPO are usually the simpler fit.
Sources
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning (Shao et al., introduces GRPO): https://arxiv.org/abs/2402.03300
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (v2, January 2026): https://arxiv.org/abs/2501.12948
- Proximal Policy Optimization Algorithms (Schulman et al.): https://arxiv.org/abs/1707.06347
- Training language models to follow instructions with human feedback (InstructGPT): https://arxiv.org/abs/2203.02155
- ZeRO (16 bytes per parameter breakdown): https://arxiv.org/abs/1910.02054
- LoRA Without Regret (Thinking Machines): https://thinkingmachines.ai/blog/lora/
- TRL GRPOTrainer docs: https://huggingface.co/docs/trl/main/en/grpo_trainer