acl-artifact-evaluation
Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-…
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluationInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of acl-artifact-evaluation
acl-artifact-evaluation scores 87/100 on our quality scale, 532nd of 1,179 Content & Media skills we index (top 46%).
Its SKILL.md is 5.7 KB long, well organised into 14 sections with 3 code examples: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so acl-artifact-evaluation is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
acl-artifact-evaluation compared with similar skills
All 4 of these similar skills score higher than acl-artifact-evaluation; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| acl-artifact-evaluation (this skill)by brycewang-stanford | 87 | 1.2k | 18d ago | SKILL.md |
| siyuanby siyuan-note | 100 | 46.6k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install acl-artifact-evaluation?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation. The install tabs above show the steps for each supported agent. - Which AI agents does acl-artifact-evaluation work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is acl-artifact-evaluation safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is acl-artifact-evaluation still maintained?
- The repository was last updated 18 days ago, so acl-artifact-evaluation is actively maintained.
Skill content
View source on GitHubname: acl-artifact-evaluation description: Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-acceptance public release.
ACL Artifact Evaluation
Use this to plan the evidence package around an ACL paper. ACL has no separate artifact-badge track; instead, artifact scrutiny is folded into review through the supplement archive and Section B ("scientific artifacts") of the Responsible NLP checklist, which reviewers cross-check against the PDF.
What counts as an artifact here
- Code: training/inference scripts, evaluation harnesses, prompt templates.
- Data: new corpora, annotations, filtered subsets of existing corpora, test suites, adversarial sets.
- Model outputs: generations, ranked lists, logits used in analysis — often the cheapest way to make an LLM paper checkable without GPUs.
- Human-subject materials: annotation guidelines, interface screenshots, consent text, compensation description.
Submission-time packaging rules
- Supplements upload as .tgz/.zip through the OpenReview form; links to tracked cloud storage are not acceptable, and any linked page must be anonymous.
- Scrub identity everywhere reviewers can look: file paths, git metadata, notebook author fields, license headers, dataset hosting pages, README contact lines.
- Reviewers are not required to open supplements. The paper plus checklist must stand alone; the archive is for verification, not for essential content.
Checklist items your artifact must satisfy
| Responsible NLP item (Section B) | Artifact implication | |---|---| | Cited creators + versions of used artifacts | Pin dataset/model versions in the README and bibliography | | License / terms of use stated | Include the license you release under and those you consumed under | | Use consistent with intended use | Justify research use of scraped or user-generated data | | PII and offensive content handled | Describe scanning/anonymization steps actually performed | | Documentation of domains, languages, demographics | Ship a data statement or datasheet, not just row counts | | Statistics on splits reported | Train/dev/test sizes in both paper and README |
Checklist answers contradicted by the archive read as misleading information — grounds for desk rejection under ARR policy, and a credibility wound even when not enforced.
What an ACL reviewer opens first
- The README — it has roughly one minute to orient them.
- Prompt files and evaluation scripts, for any LLM claim: exact prompts, decoding parameters, and scoring code are the reproduction spine.
- Annotation guidelines, for any dataset or human-eval claim: reviewers judge whether the labels could possibly mean what the paper says.
- A sample of the data itself — quality problems visible in twenty random examples have sunk otherwise strong resource papers.
Vignette: packaging a multilingual benchmark submission
A hypothetical paper releases a 7-language reading-comprehension test suite built from news text plus a baseline evaluation of five LLMs.
- Ship per-language provenance: source, license, collection window, and the filtering pipeline as runnable code, since "web text" alone fails checklist item B on documentation.
- Include annotator guidelines, pay, recruitment channel, and agreement statistics; multilingual annotation quality is the first attack surface.
- Provide the exact prompts and outputs for all five models so reviewers can re-score without API keys.
- Keep a versioned, hash-stamped test file so post-publication contamination can be audited later.
Release ladder after acceptance
anonymous supplement -> public repo + dataset page -> archived, versioned release
(review-time) (camera-ready links) (DOI/hub artifact, cited version)
Post-acceptance, register the artifact where your community actually looks (model/dataset hubs, a maintained repo), state the license explicitly, and put the citation-of-record (the Anthology entry) in the README.
Anonymization sweep, concretely
Run these before zipping, on a copy:
# authorship trails in code and docs
grep -ri "yourname\|yourlab\|university" . --include="*.py" --include="*.md"
# git history and remotes leak owners
rm -rf .git; # or re-init a fresh repo for the archive copy
# notebook metadata carries usernames and kernel paths
jupyter nbconvert --clear-output --inplace *.ipynb
# absolute paths in configs and logs
grep -r "/home/\|/Users/" . | head
Then check the parts tools miss: license headers naming the lab, dataset hosting pages with institutional branding, model cards listing maintainers, and README badges pointing at owner-named CI.
Sizing and format sanity
- Keep the archive lean: strip checkpoints reviewers cannot load anyway, cached datasets, and virtualenvs; describe big assets and provide them at camera-ready instead.
- One top-level README, one environment file, one entry point per claimed result — reviewers grant roughly a minute before giving up.
- Verify the .zip/.tgz opens on a machine that has never seen the project; OpenReview upload limits and accepted fields vary by cycle, so check the live form rather than last cycle's.
Output format
[Artifact role] anonymous supplement / camera-ready release / public benchmark
[Contents] <code/data/prompts/outputs/guidelines>
[Checklist alignment] <Section B items satisfied vs missing>
[Anonymity findings] <paths/metadata/hosting leaks>
[Release plan] <post-acceptance registry, license, versioning>
Related Skills
siyuan
46.6kAn open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
