Explainable Reinforcement Learning
The most comprehensive XRL paper list: 277 papers (2016–2026) on interpretable and explainable RL. Surveys, saliency, counterfactuals, policy summarization and more.
Install / Use
npx skills add yanzheb/explainable-reinforcement-learningInstalls into whichever agent you are using.
README
Explainable Reinforcement Learning (XRL) Resources
Explainable reinforcement learning (XRL) aims to make the goals, decisions, and learned behavior of RL agents understandable to humans. This repository is a curated, comprehensive bibliography of the field: 18 surveys, 258 articles, and 1 overview published between 2016 and 2026.
Topics include interpretable policies, saliency and attribution methods, counterfactual explanations, policy and trajectory summarization, reward decomposition, and programmatic or neuro-symbolic agents.
⭐ Star the repo to help others find it. Open an issue to suggest a paper or report an error, or reach out at contact (at) xrl (dot) ai.
How to Use This List
New to XRL? Begin with Survey Papers for a grounded overview of the field. To keep up with current work, skim Recently Added; to dive into a specific year or venue, use the Table of Contents and the per-year tables under Papers.
Table of Contents
- Resources
- Survey Papers (18 papers)
- General XRL Surveys (10) | Deep RL Explainability (4) | Focused Topics and Domains (4)
- Papers (258 papers)
Recently Added
- Aligning Agent Policies with Preferences: Human-Centered Interpretable Reinforcement Learning — AIES, 2025
- LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement Learning — ICLR, 2025
- Interpretable Deep Reinforcement Learning Via Concept-Based Policy Distillation — Mach. Learn., 2025
- Hierarchical Programmatic Option Framework — NeurIPS, 2024
- Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents — NeurIPS, 2024
- TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models — CoRR, 2025
- Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments — CoRR, 2025
- Neural DNF-MT: A Neuro-symbolic Approach for Learning Interpretable and Editable Policies — AAMAS, 2025
- Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning — ICML, 2025
- Interpreting Emergent Planning in Model-Free Reinforcement Learning — ICLR, 2025
Resources
- Awesome Explainable Reinforcement Learning — curated list of XRL papers and resources
- AAAI 2024 Workshop on eXplainable AI approaches for Deep Reinforcement Learning (XAI4DRL) — workshop proceedings on OpenReview
- RLC 2024 Workshop on Interpretable Policies in Reinforcement Learning (InterpPol) — workshop proceedings on OpenReview
- RLC 2025 Workshop on Programmatic Reinforcement Learning (PRL) — workshop proceedings on OpenReview
- Explanations for Sequential Decision-Making - an Overview — AAAI 2026 overview by Baier et al.
Survey Papers
General XRL Surveys
| #/Link | Title | Venue/Journal | Year | |:---:|:---|:---|:---:| | 1 | A Survey of Explainable Reinforcement Learning: Targets, Methods and Needs | CoRR | 2025 | | 2 | Explainable Reinforcement Learning: A Survey and Comparative Review | ACM Comput. Surv. | 2024 | | 3 | A survey on interpretable reinforcement learning | Mach. Learn. | 2024 | | 4 | Explainable reinforcement learning (XRL): a systematic literature review and taxonomy | Mach. Learn. | 2024 | | 5 | A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges | CoRR | 2023 | | 6 | Explainable reinforcement learning for broad-XAI: a conceptual framework and survey | Neural Comput. Appl. | 2023 | | 7 | Explainability in reinforcement learning: perspective and position | CoRR | 2022 | | 8 | Explainable AI and Reinforcement Learning - A Systematic Review of Current Approaches and Trends | Frontiers Artif. Intell. | 2021 | | 9 | Explainable Reinforcement Learning: A Survey | CD-MAKE | 2020 | | 10 | Reinforcement Learning Interpretation Methods: A Survey | IEEE Access | 2020 |
Deep RL Explainability
| #/Link | Title | Venue/Journal | Year | |:---:|:---|:---|:---:| | 1 | A Survey on Explainable Deep Reinforcement Learning | CoRR | 2025 | | 2 | Explainability in Deep Reinforcement Learning, a Review into Current Methods and Applications | ACM Comput. Surv. | 2023 | | 3 | Explainable Deep Reinforcement Learning: State of the Art and Challenges | ACM Comput. Surv. | 2022 | | 4 | Explainability in deep reinforcement learning | Knowl. Based Syst. | 2021 |
Focused Topics and Domains
| #/Link | Title | Venue/Journal | Year | |:---:|:---|:---|:---:| | 1 | Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities | ACM Comput. Surv. | 2024 | | 2 | A Survey of Global Explanations in Reinforcement Learning | Explainable Agency in Artificial Intelligence | 2024 | | 3 | Explainable and Interpretable Reinforcement Learning for Robotics | SLAIML | 2024 | | 4 | Advances in Explainable Reinforcement Learning: An Intelligent Transportation Systems Perspective | Explainable Artificial Intelligence for Intelligent Transportation Systems | 2023 |
Papers
Top venues: NeurIPS (19) | ICLR (16) | AAAI (11) | ICML (10) | AAMAS (9) | CoRR (9) | XAI Workshop @ IJCAI (5) | Artif. Intell. (4) | IEEE Robotics Autom. Lett. (4) | IJCAI (4) | IROS (4) | SSCI (4) | AIES (3) | ECML PKDD (3) | Eng. Appl. Artif. Intell. (3)
2025
| #/Link | Title | Venue/Journal | Year | |:---:|:---|:---|:---:| | 1 | Neural DNF-MT: A Neuro-symbolic Approach for Learning Interpretable and Editable Policies | AAMAS | 2025 | | 2 | Aligning Agent Policies with Preferences: Human-Centered Interpretable Reinforcement Learning | AIES | 2025 | | 3 | Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environment
Related Skills
mcp
Use the `mcp_perplexity-ask_perplexity_search` tools to answer questions. You should use this instead of the `web_search` tool because it is a lot more accurate.
practical-power-systems-synthesis
This skill enables synthesis in the domain of power-systems (engineering). It represents research-level-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform synthesis operations related to power-systems.
semi-supervised-optogenetics-testing
This skill enables testing in the domain of optogenetics (neuroscience). It represents intermediate-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform testing operations related to optogenetics.
data-mining-interpretation-fundamental
This skill enables interpretation in the domain of data-mining (data-science). It represents fundamental-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform interpretation operations related to data-mining.
Security Score
Audited on Aug 3, 2026
