acl-experiments
Use when designing or auditing experiments for an ACL paper, covering tuned LLM baselines, multi-dataset and multilingual evaluation, statistical significance and variance, human evaluation with agreement reporting, contamination and prompt-sensitivity controls, ablations, and error-analysis expecta…
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-experimentsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of acl-experiments
acl-experiments scores 87/100 on our quality scale, 454th of 965 AI & Machine Learning skills we index (top 48%).
Its SKILL.md is 5.6 KB long, well organised into 11 sections with 2 code examples: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so acl-experiments is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
acl-experiments compared with similar skills
All 4 of these similar skills score higher than acl-experiments; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| acl-experiments (this skill)by brycewang-stanford | 87 | 1.2k | 18d ago | SKILL.md |
| claude-memby thedotmack | 100 | 95.2k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.1k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install acl-experiments?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-experiments. The install tabs above show the steps for each supported agent. - Which AI agents does acl-experiments work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is acl-experiments safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is acl-experiments still maintained?
- The repository was last updated 18 days ago, so acl-experiments is actively maintained.
Skill content
View source on GitHubname: acl-experiments description: Use when designing or auditing experiments for an ACL paper, covering tuned LLM baselines, multi-dataset and multilingual evaluation, statistical significance and variance, human evaluation with agreement reporting, contamination and prompt-sensitivity controls, ablations, and error-analysis expectations in NLP reviewing.
ACL Experiments
Use this while the experimental story can still change. The ACL evidence bar is not "beats the baseline once": it is a defensible measurement of a language capability, with the failure modes examined.
Baseline honesty
- Include the strongest cheap baseline: a well-prompted current LLM has become mandatory context for most tasks — a method beating only pre-LLM systems invites the "does this matter now?" review.
- Tune baselines with the same care as your method (same search budget, same data); reviewers explicitly probe for asymmetric tuning.
- Report the trivial baselines (majority class, copy input, retrieval-only) when they contextualize how hard the task actually is.
Evaluation design
- Breadth must match the claim: a "general" claim needs multiple datasets; a cross-lingual claim needs typologically distinct languages, not three Romance neighbors.
- Automatic metrics need justification for generation tasks — pair n-gram or embedding metrics with human or LLM-judge evaluation, and validate any LLM-judge against human labels before leaning on it.
- Fix the evaluation protocol before final runs: dev-set peeking on the test set via repeated submissions is unreportable and unrepairable.
Statistical floor
| Result flavor | Required rigor at ACL | |---|---| | Small deltas between systems | Significance test (bootstrap/permutation) or overlapping-interval honesty | | Fine-tuning results | Multiple seeds; mean and deviation in the table, defined in the caption | | Prompted-LLM results | Multiple prompt paraphrases and/or samples; sensitivity range reported | | Human evaluation | Raters per item, agreement statistic (e.g., Krippendorff's alpha), pay disclosed | | Correlation claims (metrics) | Confidence intervals and comparison against existing metric correlations |
The Responsible NLP checklist (Section C) asks for descriptive statistics and error bars — an experiment plan that cannot fill Section C truthfully is incomplete by construction.
Contamination and validity controls
- Reason explicitly about test-set membership in pretraining data: release dates vs model cutoffs, overlap scans, or held-back fresh test items.
- Watch prompt leakage: few-shot exemplars drawn from the test distribution, instructions embedding label hints.
- For annotation-based data, quantify label quality before measuring models against it; models are now frequently better than noisy gold labels.
Ablations and the mechanism claim
- Each component the abstract credits needs an ablation row; each ablation row needs the same variance treatment as the headline number.
- Prefer ablations that test the explanation (e.g., "gains come from the retrieval step") over combinatorial component sweeps.
- Scale ablation: if a claim is "method X helps," show it at two model sizes or state the single-scale limitation explicitly.
Error analysis as a deliverable
The distinctive ACL expectation: a quantitative error analysis with named categories.
- Sample failures (100-200) from the strongest configuration.
- Induce 4-8 functional error categories; double-annotate a subset and report agreement.
- Report category frequencies for your method vs the best baseline — where do gains actually come from?
- Feed the two most persistent categories into Limitations.
Pre-run design worksheet
Claim: <one sentence>
Datasets: <n, why these, language list>
Baselines: <incl. tuned LLM baseline + trivial floor>
Runs/variance: <seeds or prompt paraphrases; interval type>
Significance: <test, when applied>
Human eval: <items, raters, agreement plan, pay>
Contamination: <audit method>
Ablations: <component -> table row>
Error analysis:<sample size, category plan>
Common evidence failures seen in ARR reviews
- Averaging over languages to hide that one language regressed — report the per-language block; reviewers open the appendix table first when a claim says "multilingual."
- Comparing your tuned method against baseline numbers copied from papers that used different preprocessing or splits.
- Treating an LLM judge as ground truth without reporting its agreement with humans on a calibration subset.
- Claiming efficiency without wall-clock, memory, or cost on matched hardware.
- Running the significance test only on the comparison that wins.
- Reporting the best seed as the headline and the mean in the appendix — reviewers call this out by name.
When compute is the constraint
- Pre-register (internally) which single configuration gets the full multi-seed treatment, and make it the headline setting.
- Use paired designs — same items, both systems — so smaller samples still yield tight comparisons and permutation tests apply cleanly.
- Prefer breadth at small scale plus depth at one large scale over a thin sweep of everything; state the choice in the setup section.
- Cache and release intermediate outputs so ablations re-score rather than re-run.
Output format
[Evidence verdict] convincing / thin / misaligned-with-claim
[Baseline gaps] <missing or under-tuned comparators>
[Statistical gaps] <variance/significance/agreement omissions>
[Validity threats] <contamination/leakage/label-quality>
[Highest-value next run] <one experiment>
Related Skills
claude-mem
95.2kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
85.1kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
