ase-experiments
Use when designing or auditing the evaluation of an ASE (IEEE/ACM Automated Software Engineering) paper, covering real subject systems, fair runnable tool baselines, task-matched effectiveness metrics, ablations that isolate a learned component, oracle and correctness validation, contamination-aware…
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill ase-experimentsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of ase-experiments
ase-experiments scores 87/100 on our quality scale, 1674th of 2,889 Automation skills we index.
Its SKILL.md is 5.2 KB long, well organised into 11 sections with 2 code examples: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 21 days ago, so ase-experiments is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-06. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
ase-experiments compared with similar skills
All 4 of these similar skills score higher than ase-experiments; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| ase-experiments (this skill)by brycewang-stanford | 87 | 1.2k | 21d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 92.1k | 20d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.9k | today | MCP Server |
| rufloby ruvnet | 100 | 74.0k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install ase-experiments?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill ase-experiments. The install tabs above show the steps for each supported agent. - Which AI agents does ase-experiments work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is ase-experiments safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is ase-experiments still maintained?
- The repository was last updated 21 days ago, so ase-experiments is actively maintained.
Skill content
View source on GitHubname: ase-experiments description: Use when designing or auditing the evaluation of an ASE (IEEE/ACM Automated Software Engineering) paper, covering real subject systems, fair runnable tool baselines, task-matched effectiveness metrics, ablations that isolate a learned component, oracle and correctness validation, contamination-aware LLM handling, and provenance for mining.
ASE Experiments
Match the evidence to the automation's claim. ASE evaluations are judged on whether a tool or technique actually does what it claims on real subjects, compared fairly against the closest runnable automation. This is the axis reviewers weight most, and the one that most often becomes a Revision criterion.
Start from the claim shape
Different automations demand different evidence:
| Automation claim | Evidence that matches | Common failure | |---|---|---| | Detection (bugs, smells, vulnerabilities) | Precision/recall/F on real defects with a defined ground truth | Synthetic-only defects; unclear ground truth | | Generation / synthesis (tests, code, patches) | Validity of the produced artifact (compiles, passes, holds the property) | Similarity-to-reference proxy instead of validity | | Repair | Verified behavior change: re-run + oracle; assertion/spec preservation | "Plausible patch" without an overfitting check | | Localization / ranking | Rank-based effectiveness on real faults vs. alternatives | Cherry-picked programs; one metric only | | Scalability / performance | Real-system sizes, wall-clock with a fair config | Toy inputs; unequal baseline budget |
Real subject systems
- Use real software — open-source projects, real bug/defect datasets, real CI logs — not toy programs you constructed to make the tool look good.
- Report subject provenance: names, versions/commit SHAs, sizes, and the extraction date. Reviewers reproduce from this.
- Justify subject selection and disclose exclusions; self-selected subjects are the classic external-validity threat.
Fair, runnable tool baselines
- Compare against the closest runnable automation, configured at an equal, documented budget (time, iterations, tuning, seeds). ASE reviewers routinely rerun or scrutinize baselines.
- Pin baseline versions/commits and note reimplementation vs. original.
- If no tool baseline exists, construct a defensible non-trivial baseline (a static rewrite, a random or heuristic variant) rather than comparing only to "nothing."
Ablations that isolate the automation
If a learned or LLM component is involved, run an ablation that removes it and keeps the rest, so the marginal value of the design is visible. This is what defeats the "the model did it, not your technique" objection and keeps the paper ASE-shaped rather than ML-shaped.
Oracles and correctness
- State the oracle explicitly: how do you know a generated test is meaningful, or a repair is correct? Re-execution, differential testing, formal checks, or human audit — name it.
- For repair/synthesis, guard against overfitting to the evaluation oracle (e.g., patches that pass the given tests but break behavior): report a held-out or manual correctness check.
Statistics and effect sizes
- Report effect sizes and dispersion (confidence intervals, non-parametric tests where appropriate), not just point estimates or a single accuracy number.
- For randomized techniques (search-based, sampling, LLM temperature > 0), report repeated runs with variance and fix/seed the randomness for the artifact.
Contamination-aware LLM handling
- Record model identifiers and dates; a model updated between runs invalidates comparisons.
- Consider training-data contamination: benchmarks the model may have seen inflate results — report on held-out or post-cutoff subjects where feasible, and say so.
- Cache raw model outputs so the artifact reproduces rather than re-samples a live API.
Mining and dataset provenance
- Pin repository SHAs, the corpus extraction date, query/filter criteria, and any labeling protocol with inter-rater agreement for manually coded data.
- Version the dataset and describe how to regenerate it; a package that needs live scraping re-samples a moving target.
Evaluation audit checklist
[Claim-evidence] each claim -> a matching metric on real subjects (not a proxy)
[Subjects] real, provenance-pinned, selection justified, exclusions disclosed
[Baselines] closest runnable tool, version pinned, equal documented budget
[Ablation] learned/LLM component isolated; marginal value of the design shown
[Oracle] correctness defined; overfitting-to-oracle checked
[Stats] effect sizes + dispersion; repeated runs for randomized methods
[LLM] model IDs/dates recorded; contamination considered; outputs cached
[Repro] provenance pinned; dataset/tool versioned for the artifact
Output format
[Automation claim] detection / generation / repair / localization / scalability
[Evidence match] metric(s) that fit the claim, on real subjects
[Baseline fairness] closest tool, budget parity, versions
[Ablation + oracle] learned-component ablation present; correctness oracle stated
[Threats] subject selection / oracle validity / baseline fairness / contamination — bounded how?
Related Skills
Agent-Reach
92.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Scrapling
85.9k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
ruflo
74.0k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
