aistats-experiments
Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than chase leaderboards.
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aistats-experimentsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of aistats-experiments
aistats-experiments scores 83/100 on our quality scale, 2899th of 4,610 Development & Engineering skills we index.
Its SKILL.md is 3.4 KB long, split into 7 sections with 1 code example: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so aistats-experiments is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
aistats-experiments compared with similar skills
All 4 of these similar skills score higher than aistats-experiments; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| aistats-experiments (this skill)by brycewang-stanford | 83 | 1.2k | 18d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.8k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 3d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
Frequently asked questions
- How do I install aistats-experiments?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aistats-experiments. The install tabs above show the steps for each supported agent. - Which AI agents does aistats-experiments work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is aistats-experiments safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is aistats-experiments still maintained?
- The repository was last updated 18 days ago, so aistats-experiments is actively maintained.
Skill content
View source on GitHubname: aistats-experiments description: Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than chase leaderboards.
AISTATS Experiments
Use this before submission when the empirical or simulation story is not yet locked.
Experiment audit
- Map each empirical claim to a table, figure, simulation, ablation, or robustness check.
- Include baselines that represent both ML practice and relevant statistical methods.
- Separate synthetic simulations that validate assumptions from real-data experiments that show practical relevance.
- Report uncertainty for stochastic results: repeated runs, standard errors, confidence intervals, paired tests, or bootstrap intervals when appropriate.
- Report dataset splits, preprocessing, metrics, hyperparameter search ranges, final chosen settings, selection criteria, random seeds, hardware, software versions, and runtime.
- Add ablations for the mechanism, not just cosmetic variants.
- Audit for leakage, selection bias, multiple-comparison issues, and mismatch between theoretical assumptions and empirical setup.
What experiments are for at this venue
- AISTATS experiments exist to validate theory, not to win leaderboards. One focused simulation confirming a predicted rate outweighs five extra benchmark datasets.
- The strongest design triad: a synthetic study where assumptions hold exactly, a study where they are deliberately violated, and a real-data study showing practical behavior.
- Reviewers, frequently statisticians, check whether the empirical regime — sample size, dimension, noise level — matches the asymptotic regime of the theorems. A bound proven as n grows but tested only at n = 500 invites the question of relevance.
Theory-validation design table
| Theoretical claim | Matching experiment | Reject pattern avoided | |---|---|---| | Convergence rate in n | Log-log error versus n with fitted slope | "Rates asserted but never plotted" | | Confidence-interval coverage | Empirical coverage across many replications | "Nominal 95 percent never verified" | | Regret bound | Cumulative regret versus horizon, with the bound curve overlaid | "Bound and trajectory never compared" | | Robustness to misspecification | Violation-severity sweep | "Guarantees hold under assumptions the experiments quietly break" |
Vignette: a kernel conditional independence test
Suppose the paper proves finite-sample type-I error control under a boundedness assumption. The matching plan: simulate under the null at several sample sizes to verify size, sweep dependence strength for power curves, then inject heavy-tailed noise that breaks boundedness to map degradation — every panel tied to a numbered theorem or remark.
Statistical reporting floor
- Replication counts and seeds for every stochastic figure; captions must say whether bars are standard errors, confidence intervals, or quantiles.
- Report the compute actually consumed rather than vague feasibility language.
Output format
[Experiment readiness] strong / adequate / weak
[Claim -> evidence map] <claim: table/figure/simulation>
[Missing statistical evidence] <uncertainty/test/seed/baseline>
[Reproducibility gaps] <hyperparameters/compute/data/code>
[Decision-critical next run] <one experiment or simulation>
Related Skills
ai-job-search
44.8kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
