acmmm-experiments
Use when designing or auditing the experiments of an ACM MM (ACM Multimedia) paper — matched baselines per modality, ablations that isolate the cross-modal fusion, user studies or QoE measurement where the claim is subjective, dataset and media licensing, and honest compute reporting, so evidence su…
Install / Use
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acmmm-experimentsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Customer SupportSupported Platforms
Our assessment of acmmm-experiments
acmmm-experiments scores 87/100 on our quality scale, 163rd of 334 Customer Support skills we index (top 49%).
Its SKILL.md is 5.2 KB long, well organised into 10 sections with 2 code examples: a solid amount of guidance for an agent.
With 1,158 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so acmmm-experiments is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
acmmm-experiments compared with similar skills
All 4 of these similar skills score higher than acmmm-experiments; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| acmmm-experiments (this skill)by brycewang-stanford | 87 | 1.2k | 18d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 10d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install acmmm-experiments?
- Run
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acmmm-experiments. The install tabs above show the steps for each supported agent. - Which AI agents does acmmm-experiments work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is acmmm-experiments safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is acmmm-experiments still maintained?
- The repository was last updated 18 days ago, so acmmm-experiments is actively maintained.
Skill content
View source on GitHubname: acmmm-experiments description: Use when designing or auditing the experiments of an ACM MM (ACM Multimedia) paper — matched baselines per modality, ablations that isolate the cross-modal fusion, user studies or QoE measurement where the claim is subjective, dataset and media licensing, and honest compute reporting, so evidence supports a multimedia claim.
ACM MM Experiments
Use this to make an ACM Multimedia paper's evidence match its claim. The reviewer's implicit questions are: does it work, does the cross-modal part cause the gain, when does it fail, and — if the target is perceptual — do people actually prefer it.
The four questions and how to answer them
| Question | Evidence that answers it | |---|---| | Does it work? | The headline metric on a recognized benchmark, against strong, matched baselines | | Does the fusion cause the gain? | A leave-one-modality-out / component ablation isolating the cross-modal term | | When does it fail? | Failure cases per modality (e.g., noisy audio, missing captions) shown honestly | | Do people prefer it? | A user study with reported N, protocol, and inter-rater agreement — for subjective claims |
The second row is what separates an ACM MM experiment section from a single-modality one: if removing a modality does not move the result, the paper is not really cross-modal.
Matched baselines
- Compare against the strongest existing method, re-run under your data and preprocessing where feasible, not a weakened reimplementation.
- Include a late-fusion / naive-concatenation baseline so the reader sees what the fancy fusion buys over the obvious one.
- Hold everything but the mechanism fixed: same backbone, same features, same training budget, so the delta is attributable.
Ablations that isolate the cross-modal claim
Full model .................... reference
- audio stream ................ tests whether audio carries signal
- text/caption stream ......... tests whether language carries signal
- alignment / fusion module ... replaced by concatenation: tests the MECHANISM
- synchronization assumption .. shuffled timing: tests whether cross-modal timing matters
Report each ablation with the same metric and variance as the headline, and state which term carries most of the gain — reviewers reward a paper that can point to why it works.
User studies and QoE
When the claim is subjective (quality, naturalness, engagement, aesthetics), a benchmark number is not enough:
- Pre-register the protocol: task, number of raters, stimuli, and the question asked.
- Report inter-rater agreement and a significance test, not just a mean preference.
- Describe compensation and consent briefly; a study a reviewer cannot assess is discounted.
Data, media, and compute honesty
- State dataset licenses and any consent/usage terms for media, especially for user-generated or scraped content.
- Report compute (hardware, training time) so cost is legible; a cross-modal model that only wins at 10x compute should say so.
- Fix and report seeds; report variance over runs where the margin is small.
Benchmarks and metrics per modality
A cross-modal paper is judged against each community's expectations at once, so pick metrics each sub-field recognizes rather than a single convenient number.
- Retrieval / recommendation — report ranking metrics (Recall@K, mAP, NDCG) and say which gallery/query split, because cross-modal retrieval numbers are split-sensitive.
- Generation / synthesis — pair a distributional metric with a human/QoE judgment; automatic scores for generated media correlate imperfectly with perceived quality.
- Recognition / detection — use the standard task metric, but show the multimodal case, not only the clean single-modality one.
- Systems / delivery — report latency, throughput, and bitrate/quality trade-offs, not just accuracy.
State the metric's direction and any threshold, and keep the same metric across the headline table and every ablation so the reader can trace the fusion's contribution row by row.
Statistical reporting
- Report variance (standard deviation or confidence interval) over runs when margins are small; a single-seed win on a close benchmark is not persuasive.
- For user studies, report a significance test and inter-rater agreement, not a bare mean.
- Do not average away modality-specific behavior: a model that helps on audio-rich clips and hurts on silent ones should show that split, not hide it in a global mean.
Common ACM MM experiment failures
- Fusion that does not matter — ablations show no modality is load-bearing.
- Weak baselines — beating only a vision-only or text-only strawman.
- Asserted perception — "more natural" with no user study.
- Unlicensed media — datasets used without stating rights.
- Hidden cost — big gains that quietly require far more compute.
Output format
[Works] strong/matched baselines / weak or unmatched: <which>
[Fusion causal] ablation isolates the mechanism / does not
[Failure analysis] present per modality / missing
[Perceptual claim] user study with agreement / asserted
[Data + compute] licenses and cost reported / gaps: <list>
[Top fixes] <ordered before submission or rebuttal>
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
