experiment-design
Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".
Install / Use
npx skills add fcakyon/phd-skills --skill experiment-designInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Education & ResearchSupported Platforms
Our assessment of experiment-design
experiment-design scores 85/100 on our quality scale, 305th of 435 Education & Research skills we index.
Its SKILL.md is 3.9 KB long, well organised into 10 sections with 2 code examples: a solid amount of guidance for an agent.
It has 406 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 18 days ago, so experiment-design is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
experiment-design compared with similar skills
All 4 of these similar skills score higher than experiment-design; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| experiment-design (this skill)by fcakyon | 85 | 406 | 18d ago | SKILL.md |
| last30days-skillby mvanhorn | 100 | 63.5k | today | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install experiment-design?
- Run
npx skills add fcakyon/phd-skills --skill experiment-design. The install tabs above show the steps for each supported agent. - Which AI agents does experiment-design work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is experiment-design safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is experiment-design still maintained?
- The repository was last updated 18 days ago, so experiment-design is actively maintained.
Skill content
View source on GitHubname: experiment-design description: > Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".
Experiment Design Methodology
You are helping a researcher design rigorous experiments. Follow this methodology systematically.
Step 1: Understand the Research Question
Before designing any experiment:
- Ask what specific hypothesis or claim the experiment should support
- Identify the dependent variable (metric) and independent variables (factors)
- Clarify the baseline: what is the current best result or default configuration?
Step 2: Single-Variable Isolation
Every ablation study must change exactly ONE variable at a time. For each factor:
- Define the factor — what is being varied (e.g., loss function, learning rate, architecture component)
- List levels — all values this factor will take (e.g., CE, focal, VAR)
- Fix everything else — document what stays constant (seed, data split, epochs, hardware)
- Predict outcome — before running, state what you expect and why
Template for each ablation row:
| Run ID | Factor | Value | Fixed Config | Expected Outcome |
|--------|--------|-------|-------------|-----------------|
Step 3: Experiment Matrix
For multi-factor studies, use a structured matrix:
- Full factorial — if factors are few (≤3) and levels are few (≤3 each)
- Sequential elimination — if factors are many: run single-factor ablations first, then combine winners
- Latin square — if full factorial is too expensive: sample representative combinations
Always calculate total runs before committing:
Total runs = product of all factor levels
GPU hours = total runs × hours_per_run
Step 4: Resource Estimation
For each experiment plan, estimate:
- GPU hours: runs × time_per_run (check with user's hardware)
- API costs: if using external APIs (Gemini, OpenAI), estimate tokens × price
- Wall clock time: accounting for sequential dependencies and GPU availability
- Storage: checkpoint sizes × number of runs
Flag if total cost exceeds reasonable bounds and suggest prioritization.
Step 5: Config Stub Generation
Generate configuration stubs that match the user's existing config format. Read existing configs first to match:
- File format (YAML, JSON, TOML)
- Key naming conventions
- Directory structure for outputs
- Logging/tracking integration (wandb, neptune, tensorboard)
Step 6: Execution Plan
Create a concrete execution plan:
- Order runs by dependency (baselines first, then ablations)
- Identify which runs can be parallelized across GPUs
- Create a shell script or batch runner matching the project's existing patterns
- Include checkpointing strategy for long runs
Step 7: Analysis Plan
Before running, define how results will be analyzed:
- Which metrics to compare (primary + secondary)
- Statistical significance test if applicable (paired t-test, bootstrap CI)
- How to handle failed/crashed runs
- Visualization: what plots to generate (comparison tables, bar charts, learning curves)
Verification Checkpoints
Before finalizing the experiment plan:
- [ ] Each ablation changes exactly one variable
- [ ] Baseline is clearly defined and will be run with same setup
- [ ] Resource estimate is within budget
- [ ] Config stubs match existing project format
- [ ] Analysis plan is defined before execution begins
- [ ] Seeds are fixed for reproducibility
Output Format
Always produce:
- Experiment matrix table — all runs with their configurations
- Resource estimate — GPU hours, API costs, storage
- Execution script — ready-to-run commands matching project conventions
- Analysis plan — metrics, comparisons, visualizations
Related Skills
last30days-skill
63.5kAI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
