grpo
Reference for the GRPO (Group Relative Policy Optimization) algorithm
Install / Use
npx skills add benchflow-ai/skillsbench --skill grpoInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of grpo
grpo scores 86/100 on our quality scale, 1412th of 2,659 Automation skills we index.
Its SKILL.md is 4.3 KB long, well organised into 8 sections with 4 code examples: a solid amount of guidance for an agent.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so grpo is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
grpo compared with similar skills
All 4 of these similar skills score higher than grpo; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| grpo (this skill)by benchflow-ai | 86 | 1.8k | 2mo ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.4k | 15d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.6k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.6k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 7d ago | SKILL.md |
Frequently asked questions
- How do I install grpo?
- Run
npx skills add benchflow-ai/skillsbench --skill grpo. The install tabs above show the steps for each supported agent. - Which AI agents does grpo work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is grpo safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is grpo still maintained?
- The repository was last updated about 2 months ago, so grpo is actively maintained.
Skill content
View source on GitHubname: grpo description: Reference for the GRPO (Group Relative Policy Optimization) algorithm. Use when implementing, debugging, or verifying a GRPO training pipeline — covers the mathematical formulation (group-relative advantages, clipped surrogate loss, KL penalty), the training loop (generate → score → advantage → loss), log-probability computation, advantage estimation, and relationship to PPO/REINFORCE.
GRPO Algorithm Reference
Overview
Group Relative Policy Optimization (GRPO) is a policy gradient method that eliminates the need for a learned critic by computing advantages from group statistics. For each prompt, multiple completions are sampled and their rewards are normalized within the group to produce advantages.
Key insight: Instead of training a value function to estimate V(s), GRPO uses the mean reward of the group as the baseline. This removes the critic entirely, reducing memory and avoiding value function approximation errors.
Training Loop
Each step follows this pipeline:
1. Sample G completions per prompt from current policy
2. Decode completions to text
3. Score completions with reward function
4. Compute group-relative advantages
5. Compute per-token log-probs (current policy + reference policy)
6. Compute clipped surrogate loss with KL penalty
7. Backpropagate and update
Advantage Estimation
Rewards are normalized within each group of G completions for the same prompt:
mu = mean(r_1, ..., r_G)
sigma = std(r_1, ..., r_G)
A_i = (r_i - mu) / (sigma + epsilon)
epsilon is a small constant (typically 1e-4 to 1e-8) for numerical stability only. It prevents division by zero when all rewards in a group are identical.
Properties:
- Advantages within each group sum to approximately zero;
- High-reward completions get positive advantages, low-reward get negative;
- The learning signal vanishes if epsilon is too large (dominates denominator) or if rewards are constant. For example, if all rewards in a group are identical (or nearly identical), then causing all advantages to collapse to 0;
Log-Probability Computation
Per-token log probabilities via log-softmax:
log_prob(token_i) = logit(token_i) - logsumexp(logits)
Critical invariant: log_prob <= 0 always, since it's the log of a probability in (0, 1].
Sequence-level log probability: log_pi(y|x) = sum_t log_pi(t_j | x, t_<j)
Loss Function
The GRPO loss combines a clipped surrogate objective with a KL divergence penalty:
ratio = exp(log_pi_theta(y|x) - log_pi_old(y|x))
L_clip = min(ratio * A, clip(ratio, 1-eps, 1+eps) * A)
L_kl = beta * KL(pi_theta || pi_ref)
loss = -E[L_clip] + L_kl
| Component | Purpose |
|-----------|---------|
| ratio | How much the policy has changed from the generation policy |
| Clipping | Prevents destructively large policy updates |
| beta * KL | Keeps the policy close to the reference (prevents degeneration) |
Key Hyperparameters
| Parameter | Typical Range | Effect |
|-----------|--------------|--------|
| num_generations (G) | 4–16 | More = lower variance advantages, higher compute cost |
| beta | 0.01–0.1 | Higher = more conservative updates (closer to reference) |
| epsilon (clip) | 0.1–0.2 | Narrower = more conservative updates |
| epsilon (advantage) | 1e-8–1e-4 | Must be small; only for numerical stability |
| learning_rate | 1e-7–5e-6 | Much lower than SFT; RL is sensitive to LR |
Available References
| File | Contents | When to load |
|------|---------------------------------------------------------------------------------------------------------------------------------------------------|-------------|
| references/grpo-algorithm.md | Full mathematical formulation with notation table, step-by-step derivations, comparison to PPO and REINFORCE | When you need to verify whether a specific implementation detail matches the algorithm specification |
| references/grpo-trainer-internals.md | GRPOTrainer implementation from popular frameworks (TRL): method-by-method breakdown, data flow diagram, and how each algorithm step maps to code | When tracing bugs through a GRPOTrainer implementation or understanding how the algorithm maps to specific code paths |
Related Skills
Agent-Reach
86.4kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.6k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
84.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
