workflows:review
Run multi-agent econometric review on estimation code, identification arguments, and research artifacts
Install / Use
npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill workflows-reviewInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of workflows:review
workflows:review scores 92/100 on our quality scale, 507th of 1,657 Automation skills we index (top 31%).
Its SKILL.md is 13 KB long, well organised into 20 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
With 4,360 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 3 days ago, so workflows:review is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-27. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
workflows:review compared with similar skills
All 4 of these similar skills score higher than workflows:review; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| workflows:review (this skill)by brycewang-stanford | 92 | 4.4k | 3d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 85.6k | 11d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 83.9k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 4d ago | SKILL.md |
Frequently asked questions
- How do I install workflows:review?
- Run
npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill "workflows:review". The install tabs above show the steps for each supported agent. - Which AI agents does workflows:review work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is workflows:review safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is workflows:review still maintained?
- The repository was last updated 3 days ago, so workflows:review is actively maintained.
Skill content
View source on GitHubname: workflows:review description: Run multi-agent econometric review on estimation code, identification arguments, and research artifacts argument-hint: "<file paths, directory, plan reference, PR number, or empty for auto-detect>" allowed-tools: Read, Grep, Glob, Bash
Review Command
Pipeline mode: This command operates fully autonomously. All decisions are made automatically.
Perform exhaustive econometric and methodological review using multi-agent parallel analysis. Domain-specific reviewers check estimation quality, identification strategy, numerical stability, and mathematical rigor.
Input
<review_target> #$ARGUMENTS </review_target>
Execution Workflow
Phase 1: Scope Detection
-
Eligibility Check
Before launching review agents, verify there is something to review. If no research artifacts are found (no estimation code, no proofs, no pipeline files, no data scripts, no output files), state "No research artifacts found to review" and stop. Do not launch agents against an empty target.
-
Determine Review Target
The review is artifact-centric: it reviews research files (estimation code, proofs, pipelines, data scripts), not git metadata. Determine the target in priority order:
- File paths or directories (e.g.,
estimation.py,src/models/,proof.tex) → review those artifacts directly - Plan reference (e.g.,
plan-3) → find the plan indocs/plans/, review files it references - PR number → fetch file list with
gh pr view --json files - Empty → auto-detect: scan the project for estimation code, proofs, pipeline files, and data scripts. If git shows recent changes, include those.
- File paths or directories (e.g.,
-
Classify Artifacts
Scan the target files and classify by type:
estimation_code: *.py with statsmodels/scipy.optimize/pyblp/linearmodels imports *.R with fixest/lfe/AER/gmm imports *.jl with Optim/NLsolve imports *.do with reg/ivregress/gmm commands simulation_code: Monte Carlo loops, DGP code, bias/RMSE computation proofs: *.tex with theorem/proof environments, *.md with derivation sections pipeline_files: Makefile, Snakefile, dvc.yaml, master.do data_code: data loading, cleaning, merge operations output_files: tables/*, figures/*, *.csv result filesThis classification drives which domain reviewers to launch.
-
Load Review Settings
Read
compound-science.local.mdin the project root. If found, usereview_agentsfrom YAML frontmatter. If the markdown body contains review context (e.g., "focus on identification strategy" or "this is a replication package"), pass it to each agent as additional instructions.If no settings file exists, use defaults:
review_agents: - econometric-reviewer - numerical-auditor - identification-critic
Protected Artifacts
The following paths are compound-science pipeline artifacts and must never be flagged for deletion or removal by any review agent:
docs/plans/*.md— Plan files created by/workflows:plandocs/brainstorms/*.md— Brainstorm files created by/workflows:brainstormdocs/solutions/*.md— Solution documents created by/workflows:compounddocs/simulations/*.md— Simulation study documentation
If a review agent flags any file in these directories for cleanup or removal, discard that finding during synthesis.
Phase 2: Agent Dispatch
Entry condition: Phase 1 classified at least one artifact; review settings loaded. Exit condition: All dispatched agents have returned findings.
Launch domain reviewers in parallel using the Task tool. The specific agents depend on artifact classification from Phase 1.
Always Run (Core Domain Review)
<parallel_tasks>
Launch econometric-reviewer, numerical-auditor, and identification-critic in parallel:
Task econometric-reviewer(changed files + review context)
→ Checks: identification strategy, endogeneity, standard errors, instrument validity,
sample selection, asymptotic properties, correct package usage
Task numerical-auditor(changed files + review context)
→ Checks: floating-point stability, convergence diagnostics, integration accuracy,
RNG seeding, matrix conditioning, overflow/underflow, gradient accuracy
Task identification-critic(changed files + review context)
→ Checks: completeness of identification argument, exclusion restriction plausibility,
functional form assumptions, parametric vs nonparametric claims, support conditions,
point vs set identification
</parallel_tasks>
Conditional Agents (Run Based on Artifact Types)
<conditional_agents>
WRITTEN ARTIFACTS: If PR contains proofs, derivations, or paper sections:
(Files matching: *.tex, *.md with theorem/proof/lemma/proposition content, docs/proofs/*)
Task journal-referee(written artifact files + review context)
→ Simulates top-5 journal referee: contribution clarity, relation to literature,
identification concerns, economic vs statistical significance, R&R concerns
(robustness, external validity, mechanism)
PIPELINE/DATA CODE: If PR contains pipeline files or data processing:
(Files matching: Makefile, Snakefile, dvc.yaml, *.do, data loading/cleaning code)
Task reproducibility-auditor(pipeline files + review context)
→ Checks: intermediate files generated by code (no manual steps), seeds documented,
package versions pinned, end-to-end pipeline, relative paths, data not committed
TABLES/FIGURES: If tables or figures were generated:
(Files matching: tables/*, figures/*, *.tex with tabular content, *.csv result files)
Task econometric-reviewer(output files + estimation code + review context)
→ Checks: table numbers match underlying code output, no manual edits to generated tables,
statistical summaries consistent with estimation logs, formatting correct
</conditional_agents>
Always Run Post-Review
Search docs/solutions/ for past issues related to this PR's modules and patterns
→ Flag matches as "Known Pattern" with links to solution docs
→ See workflows-compound/references/solution-schema.md for category detection and search workflow
Phase 3: Finding Assembly
Wait for all Phase 2 agents to complete before proceeding.
-
Collect All Findings
Gather outputs from all parallel agents into a unified findings list.
-
Categorize by Severity
| Severity | Criteria | Action | |----------|----------|--------| | CRITICAL | Incorrect identification argument, biased estimator, wrong standard errors, numerical instability producing wrong results, missing convergence check | Must fix before proceeding | | WARNING | Suboptimal estimation approach, missing robustness check, incomplete diagnostics, weak instruments not flagged, reproducibility gap | Should fix | | NOTE | Style improvements, alternative approaches worth considering, minor efficiency gains, documentation gaps | Nice to have |
-
Deduplicate and Cross-Reference
- Remove duplicate findings across agents (e.g., econometric-reviewer and identification-critic may both flag the same exclusion restriction)
- Surface solution search results: if past solutions are relevant, tag findings as "Known Pattern — see docs/solutions/[path]"
- Discard any findings that recommend deleting files in protected artifact directories
-
Estimation-Specific Synthesis
For estimation code changes, synthesize a unified assessment:
| Dimension | Status | Details | |-----------|--------|---------| | Identification | [valid/concerns/invalid] | Summary from econometric-reviewer + identification-critic | | Estimation | [correct/issues/incorrect] | Summary from econometric-reviewer + numerical-auditor | | Inference | [valid/concerns/invalid] | Standard error assessment from econometric-reviewer | | Numerical Stability | [stable/warnings/unstable] | Summary from numerical-auditor | | Reproducibility | [complete/gaps/missing] | Summary from reproducibility-auditor (if run) | | Rigor | [publication-ready/needs-work/insufficient] | Summary from journal-referee (if run) |
Phase 4: Action
-
Create Todos for All Findings
Use TodoWrite to create actionable items for all CRITICAL and WARNING findings:
TodoWrite([ { id: "review-001", task: "[CRITICAL] description", status: "pending" }, { id: "review-002", task: "[WARNING] description", status: "pending" }, ... ])For NOTES: include as a summary list — do not create individual todos unless the note is actionable.
-
Generate Review Summary
## Econometric Review Complete **Review Target:** [files/directory/plan reviewed] ### Estimation Assessment | Dimension | Status | |-----------|--------| | Identification | [status] | | Estimation | [status] | | Inference | [status] | | Numerical Stability | [status] | | Reproducibility | [status] | | Rigor | [status] | ### Findings Summary - **CRITICAL:** [count] — must fix before proceeding - **WARNING:** [count] — should fix - **NOTE:** [count] — suggestions ### CRITICAL Findings 1. [finding with agent source and file location] 2. ... ### WARNING Findings 1. [finding with agent source and file location] 2. ... ### Notes - [summarized notes] ### Known Patterns (from docs/solutions/) - [any matches from solution search] ### Review Agents Used - econometric-reviewer - numerical-auditor - identification-critic - [conditional agents if triggered] - docs/solutions/ search ### Next Steps 1. Address CRITICAL findings (must fix before proceeding) 2. Address WARNING findings (recommended) 3. Run `/workflows:compound` to document any novel solutions
Review Perspectives
The review evaluates changes from multiple research-relevant angles:
Methodological Rigor
- Is the identification strategy valid and complete?
- Are the maintained assumptions stated and plausible?
- Does the estimation approach match the identification argument?
- Are diagnostics and specification tests appropriate?
Numerical Quality
- Does estimation code handle floating-point correctly?
- Are convergence criteria appropriate?
- Is the code robust to ill-conditioned data?
- Are random seeds set for all stochastic operations?
Reproducibility
- Can results be reproduced from the replication package?
- Are all dependencies pinned?
- Does the pipeline run end-to-end without manual steps?
- Are data sources documented and accessible?
Contribution (Referee Perspective, if triggered)
- Is the contribution clearly stated?
- How does this relate to existing literature?
- Are results economically meaningful (not just statistically significant)?
- What would a skeptical referee ask for?
Configuring Review Agents
Review agents are configured in compound-science.local.md at the project root. The YAML frontmatter controls which agents run:
---
review_agents:
- econometric-reviewer
- numerical-auditor
- identification-critic
# Uncomment to always include:
# - journal-referee
# - reproducibility-auditor
---
The markdown body provides additional context passed to all review agents:
## Review Context
Focus on identification strategy — this paper uses a shift-share instrument
and we need to verify the exclusion restriction argument is complete.
To create or modify settings, edit compound-science.local.md directly.
Severity Scale
All review findings must use this severity scale:
| Severity | Meaning | Action Required | Research Example | |----------|---------|----------------|-----------------| | P0 | Invalidates core result | Must fix before proceeding | Identifi
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.3k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
83.9k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
