SkillAgentSearch skills...

aejpol-robustness

Use when an AEJ: Economic Policy manuscript's headline policy estimate needs to be shown stable and credible against specification, sample, inference, and identification threats.

Install / Use

npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aejpol-robustness

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

88/100

Supported Platforms

Universal

Our assessment of aejpol-robustness

aejpol-robustness scores 88/100 on our quality scale, 1577th of 4,610 Development & Engineering skills we index (top 35%).

Its SKILL.md is 6.4 KB long, well organised into 11 sections with 1 code example: a thorough specification that gives an agent plenty to work with.

With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
29/30
Structure
17/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 18 days ago, so aejpol-robustness is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

aejpol-robustness compared with similar skills

All 4 of these similar skills score higher than aejpol-robustness; compare them before choosing.

SkillScoreStarsUpdatedFormat
aejpol-robustness (this skill)by brycewang-stanford881.2k18d agoSKILL.md
ai-job-searchby MadsLorentzen10044.8ktodayCLAUDE.md
claude-howtoby luongnv8910041.7k3d agoCLAUDE.md
algorithmic-artby anthropics100177.9k10d agoSKILL.md
pptxby anthropics100177.9k10d agoSKILL.md

Frequently asked questions

How do I install aejpol-robustness?
Run npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aejpol-robustness. The install tabs above show the steps for each supported agent.
Which AI agents does aejpol-robustness work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is aejpol-robustness safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is aejpol-robustness still maintained?
The repository was last updated 18 days ago, so aejpol-robustness is actively maintained.

name: aejpol-robustness description: Use when an AEJ: Economic Policy manuscript's headline policy estimate needs to be shown stable and credible against specification, sample, inference, and identification threats. Organizes the robustness program by threat-to-the-policy-conclusion; it does not design the primary identification or write exhibits.

Robustness — Defending the Policy Estimate (aejpol-robustness)

When to trigger

  • The headline causal estimate moves across specifications, or you do not yet know if it does
  • A referee will ask "is this robust?" and you have no organized answer
  • Inference (clustering, few clusters, multiple outcomes) is not yet airtight
  • You need to show the policy conclusion, not just a coefficient, survives stress

Principle: robustness defends the policy conclusion, not the coefficient

At AEJ: Policy, robustness is judged by whether the policy takeaway is stable — if the headline estimate is the cost-per-job or the MVPF, show that number is stable, with its uncertainty, not merely that a regression coefficient stays significant. Organize the robustness program around the threats that would change the policy conclusion, and report enough that a skeptical referee can see each threat addressed.

Robustness by threat (each maps to a concrete check)

| Threat to the policy conclusion | Check | |---|---| | Functional form / controls drive the result | Specification ladder; show the estimate across a coherent set, not a single lucky spec | | Pre-trends / parallel-trends violation | Honest-DID (Rambachan–Roth) sensitivity bounds; placebo pre-period "effects" | | Estimator bias under staggered timing | Re-estimate with ≥1 heterogeneity-robust DID estimator (CS / SA / BJS / dCDH) | | Bandwidth / kernel (RDD) | Bandwidth sweep + bias-corrected CIs; donut-RDD if heaping at the cutoff | | Weak / invalid instrument | Effective F; AR-robust CI; over-ID test if available | | Wrong inference / few clusters | Wild-cluster bootstrap; report clustering level sensitivity | | Multiple outcomes / specifications | Romano–Wolf / sharpened q-values; a specification curve where many specs are run | | Confounding by an omitted policy/shock | Controls for co-timed policies; event-study around the focal reform only | | Selection on unobservables | Oster (2019) δ / bounds; argue the implied selection is implausible | | Sample composition / outliers | Drop influential jurisdictions; winsorize; alternative sample windows |

Sensitivity that is policy-specific

  • If the policy lesson depends on a welfare parameter you calibrate (discount rate, value of a statistic, recycling rule), report the lesson across a plausible range of that parameter, not one value.
  • If external validity is the policy worry, show heterogeneity by jurisdiction characteristics and discuss which settings the estimate travels to.

Execution bridge (StatsPAI / Stata MCP)

Run the battery, don't just enumerate it. Full map: execution-with-mcp. AEJ: Policy evaluates programs and reforms; the design must carry a policy-relevant magnitude, not just statistical significance.

  • Many outcomes / specifications: romano_wolf (step-down FWER, accounts for cross-test correlation) or benjamini_hochberg — report the adjusted threshold.
  • OVB sensitivity: oster_delta / sensemakr — the confounder strength that would overturn the headline.
  • Inference: wild_cluster_bootstrap (few clusters), twoway_cluster / conley.
  • Re-fit off one handle: audit_result(result_id) lists the missing checks and the exact suggest_function for each — no guessing the battery.
  • Exhibits: etable / did_summary_to_latex from the handle — no retyped numbers.

Keep the decisive checks in the body and the exhaustive (now actually-run) battery in the appendix. See the executed chain in the JF execution walkthrough.

Checklist

  • [ ] The headline policy number (not just a coefficient) is shown stable across specs
  • [ ] The single most likely referee threat is pre-empted with a dedicated exhibit
  • [ ] At least one heterogeneity-robust estimator shown where staggered timing applies
  • [ ] Inference stress-tested (wild-cluster / AR / multiple-testing as relevant)
  • [ ] Selection-on-unobservables addressed (Oster bounds or equivalent)
  • [ ] Calibrated welfare parameters varied across a defended range
  • [ ] No "kitchen-sink" robustness with no narrative — each check answers a named threat

Anti-patterns

  • A robustness section that is a wall of tables with no statement of which threat each rebuts
  • Showing the coefficient is stable while the welfare/policy number is never re-derived
  • A specification curve run but only the favorable region discussed
  • Treating "still significant" as robustness while ignoring magnitude stability
  • Calibrating one welfare parameter value and never probing it

Sequencing the robustness section for a referee

Order the section so a referee meets the answer before the doubt: (1) the main heterogeneity-robust estimate and its event-study; (2) the single most likely fatal threat with its dedicated check; (3) the inference stress-tests; (4) a compact specification curve or table of remaining variants; (5) the calibrated-parameter sensitivity for the welfare number. Each subsection ends with one sentence stating that the policy conclusion is unchanged, with its band — not merely that the coefficient stays signed.

Worked vignette (illustrative)

A staggered-DID estimate of a minimum-wage change on employment is the basis for a "small disemployment cost" policy claim. A referee will doubt staggered TWFE and pre-trends. The robustness program: CS and SA estimators (estimate within 10% of TWFE, illustrative), flat pre-period leads, an honest-DID bound showing the sign survives a pre-trend twice the largest observed lead, and wild-cluster inference across 30 states. The policy claim — disemployment cost per dollar of raised earnings — is re-derived under each and reported with its band.

Output format

【Headline policy number】the quantity whose stability you defend
【Top 3 threats】ranked by how badly each would change the conclusion
【Checks per threat】[threat → check → result]
【Inference】clustering / few-cluster / multiple-testing handling
【Calibrated-parameter sensitivity】range probed + conclusion stability
【Next step】aejpol-tables-figures

Related Skills

View on GitHub
GitHub Stars1.2k
CategoryDevelopment
Updated18d ago
Forks153

Languages

Stata

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions