aamas-experiments
Use when designing or auditing AAMAS experiments - self-play and population-based training, opponent selection, equilibrium and regret metrics, game-theoretic simulations, ablations, seeds, hyperparameters, compute, and claim-to-evidence fit - with emphasis on experiments that probe the interaction…
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aamas-experimentsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OtherSupported Platforms
Our assessment of aamas-experiments
aamas-experiments scores 83/100 on our quality scale, 147th of 216 Other skills we index.
Its SKILL.md is 3.8 KB long, split into 7 sections with 1 code example: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so aamas-experiments is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
aamas-experiments compared with similar skills
All 4 of these similar skills score higher than aamas-experiments; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| aamas-experiments (this skill)by brycewang-stanford | 83 | 1.2k | 18d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install aamas-experiments?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aamas-experiments. The install tabs above show the steps for each supported agent. - Which AI agents does aamas-experiments work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is aamas-experiments safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is aamas-experiments still maintained?
- The repository was last updated 18 days ago, so aamas-experiments is actively maintained.
Skill content
View source on GitHubname: aamas-experiments description: Use when designing or auditing AAMAS experiments - self-play and population-based training, opponent selection, equilibrium and regret metrics, game-theoretic simulations, ablations, seeds, hyperparameters, compute, and claim-to-evidence fit - with emphasis on experiments that probe the interaction rather than chase a single-agent leaderboard.
AAMAS Experiments
Use this before submission when the empirical or simulation story is not yet locked. At AAMAS the experiment exists to test the interaction claim, not to top a benchmark.
Experiment audit
- Map each empirical claim to a game, a self-play run, a population sweep, an ablation, or a deviation test.
- Choose opponents deliberately: self-play alone rarely suffices; include held-out opponents, population sets, or classical strategies as the claim requires.
- Separate simulations that validate a solution concept (where the equilibrium is known) from real or applied studies that show practical multiagent behavior.
- Report uncertainty for stochastic results over both seeds and opponents: standard errors, confidence intervals, or paired tests.
- Report the environment, number of agents, training regime, evaluation protocol, metrics, hyperparameter ranges, chosen settings, seeds, hardware, software versions, and runtime.
- Add ablations for the interaction mechanism (communication, reward sharing, the payment rule), not just cosmetic variants.
- Audit for the mismatch between the strategic claim and the setup: an equilibrium claim tested against only one fixed opponent, or a cooperation claim that hides a reward-shaping constant.
What experiments are for at this venue
- The strongest design shows the interaction under stress: agents that can deviate, opponents the method did not train against, and populations that vary in size or composition.
- One experiment that lets agents try to exploit the mechanism and fails to profit is worth more than five extra environments where nothing strategic is tested.
- Reviewers, often game theorists, check whether the metric matches the claim: convergence to a named solution concept, exploitability, social welfare, or regret - not just episodic return.
Interaction-validation design table
| Interaction claim | Matching experiment | Reject pattern avoided | |---|---|---| | Converges to equilibrium | Convergence/exploitability curve under simultaneous adaptation | "Equilibrium asserted, never measured" | | Mechanism is truthful | Strategic-deviation test: an agent tries to misreport | "Truthfulness proved, never stress-tested" | | Beats other agents | Round-robin vs held-out opponents and a population | "Self-play only" | | Emergent cooperation | Sweep over reward/opponent settings with variance | "One seed, one setting, one story" |
Vignette: a coordination-protocol study
Suppose the paper claims a learned protocol raises cooperation in a repeated public-goods game. The matching plan: sweep group size and defector fraction for cooperation curves, add held-out opponents that never appeared in training, and inject a free-rider agent to measure whether it profits - every panel tied to a numbered claim or definition.
Statistical reporting floor
- Seeds and replication counts for every stochastic curve; captions must state whether bands are standard errors, confidence intervals, or quantiles, and how many opponents were averaged.
- Report the compute actually consumed by self-play, not vague feasibility language.
Output format
[Experiment readiness] strong / adequate / weak
[Claim -> evidence map] <claim: game / self-play / population / deviation test>
[Missing interaction evidence] <opponents / deviation test / seeds / metric>
[Reproducibility gaps] <hyperparameters / compute / env / seeds>
[Decision-critical next run] <one experiment or simulation>
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
