trl
Reference for the TRL (Transformer Reinforcement Learning) library codebase. Use proactively before reading or editing any file under `trl/` so you have the intended contracts and invariants in mind, not just what the current code says.
Install / Use
npx skills add benchflow-ai/skillsbench --skill trlInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of trl
trl scores 88/100 on our quality scale, 393rd of 970 AI & Machine Learning skills we index (top 41%).
Its SKILL.md is 3.9 KB long, well organised into 9 sections with 2 code examples: a solid amount of guidance for an agent.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so trl is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
trl compared with similar skills
All 4 of these similar skills score higher than trl; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| trl (this skill)by benchflow-ai | 88 | 1.8k | 2mo ago | SKILL.md |
| claude-memby thedotmack | 100 | 95.2k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.0k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install trl?
- Run
npx skills add benchflow-ai/skillsbench --skill trl. The install tabs above show the steps for each supported agent. - Which AI agents does trl work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is trl safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is trl still maintained?
- The repository was last updated about 2 months ago, so trl is actively maintained.
Skill content
View source on GitHubname: trl
description: Reference for the TRL (Transformer Reinforcement Learning) library codebase. Use proactively before reading or editing any file under trl/ so you have the intended contracts and invariants in mind, not just what the current code says. Covers trainer hierarchy (SFT, DPO, GRPO, KTO), shared utility functions (selective_log_softmax, decode_and_strip_padding, padding helpers), configuration system, model wrappers, and how data flows through any TRL trainer.
TRL Library Reference
Package Structure
TRL is organized around a trainer hierarchy that extends Hugging Face transformers.Trainer.
trl/
├── trainer/
│ ├── grpo_trainer.py # GRPOTrainer
│ ├── grpo_config.py # GRPOConfig
│ ├── sft_trainer.py # SFTTrainer (supervised fine-tuning)
│ ├── dpo_trainer.py # DPOTrainer (direct preference optimization)
│ ├── kto_trainer.py # KTOTrainer (Kahneman-Tversky optimization)
│ ├── online_dpo_trainer.py # OnlineDPOTrainer
│ ├── utils.py # Shared utilities (log probs, decoding, padding)
│ └── ...
├── models/
│ └── modeling_value_head.py # Value head for PPO-style trainers
├── data_utils.py
├── commands/ # CLI entry points
└── ...
Trainer Hierarchy
All TRL trainers extend transformers.Trainer:
transformers.Trainer
├── SFTTrainer # Supervised fine-tuning
├── DPOTrainer # Direct preference optimization
├── GRPOTrainer # Group relative policy optimization
├── KTOTrainer # Kahneman-Tversky optimization
└── OnlineDPOTrainer # Online DPO
Each trainer overrides compute_loss with its specific objective, and RL-based trainers (GRPO, OnlineDPO) additionally override training_step to add a generation phase before the optimization step.
Shared Utility Functions (trainer/utils.py)
These utilities are used across multiple trainers. Read the source before modifying; the contracts below are what callers rely on.
selective_log_softmax(logits, index)
Memory-efficient per-token log-probability. Equivalent in value to F.log_softmax(logits, dim=-1).gather(...) at the selected token positions, but avoids materializing the full vocab-sized tensor.
Contract:
- Input:
logits [B, T, V],index [B, T] - Output:
log_probs [B, T], each entry a valid log-probability (i.e. non-positive) - Must agree with
F.log_softmaxto within numerical tolerance on the same inputs
decode_and_strip_padding(input_ids, tokenizer)
Converts a batch of token ID tensors into the cleaned text strings that the reward function will score.
Contract:
- Input:
input_ids [B, T], tokenizer - Output:
list[str]of lengthB - Strips padding and decoder artefacts
- Handles any reasoning-block conventions the library supports; the exact policy for complete, incomplete, and absent reasoning markers is defined in the implementation
Other Utilities
pad/pad_to_length— Pad tensors to equal or specific lengths- Various tokenizer helpers for batch processing
Configuration System
All TRL configs extend transformers.TrainingArguments. Each trainer adds its own fields:
| Config | Trainer | Key fields |
|--------|---------|------------|
| SFTConfig | SFTTrainer | max_seq_length, packing, dataset_text_field |
| DPOConfig | DPOTrainer | beta, loss_type, reference_free |
| GRPOConfig | GRPOTrainer | num_generations, beta, epsilon, reward_functions |
| KTOConfig | KTOTrainer | beta, desirable_weight, undesirable_weight |
Available References
| File | Contents | When to load |
|------|----------|-------------|
| references/trl-codebase.md | Module-by-module guide to TRL source: detailed breakdown of each trainer, model wrappers, data utilities, and CLI commands | When navigating unfamiliar parts of TRL beyond the trainer layer, or when you need details about a specific non-GRPO trainer |
Related Skills
claude-mem
95.2kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
85.0kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
