SkillAgentSearch skills...

experiment-design

Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".

Install / Use

npx skills add fcakyon/phd-skills --skill experiment-design

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

85/100

Supported Platforms

Universal

Our assessment of experiment-design

experiment-design scores 85/100 on our quality scale, 305th of 435 Education & Research skills we index.

Its SKILL.md is 3.9 KB long, well organised into 10 sections with 2 code examples: a solid amount of guidance for an agent.

It has 406 GitHub stars, a meaningful sign that others use it.

Substance
26/30
Structure
18/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 18 days ago, so experiment-design is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

experiment-design compared with similar skills

All 4 of these similar skills score higher than experiment-design; compare them before choosing.

SkillScoreStarsUpdatedFormat
experiment-design (this skill)by fcakyon8540618d agoSKILL.md
last30days-skillby mvanhorn10063.5ktodayCLAUDE.md
algorithmic-artby anthropics100177.9k12d agoSKILL.md
pptxby anthropics100177.9k12d agoSKILL.md
designby nextlevelbuilder100130.2k13d agoSKILL.md

Frequently asked questions

How do I install experiment-design?
Run npx skills add fcakyon/phd-skills --skill experiment-design. The install tabs above show the steps for each supported agent.
Which AI agents does experiment-design work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is experiment-design safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is experiment-design still maintained?
The repository was last updated 18 days ago, so experiment-design is actively maintained.

name: experiment-design description: > Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".

Experiment Design Methodology

You are helping a researcher design rigorous experiments. Follow this methodology systematically.

Step 1: Understand the Research Question

Before designing any experiment:

  • Ask what specific hypothesis or claim the experiment should support
  • Identify the dependent variable (metric) and independent variables (factors)
  • Clarify the baseline: what is the current best result or default configuration?

Step 2: Single-Variable Isolation

Every ablation study must change exactly ONE variable at a time. For each factor:

  1. Define the factor — what is being varied (e.g., loss function, learning rate, architecture component)
  2. List levels — all values this factor will take (e.g., CE, focal, VAR)
  3. Fix everything else — document what stays constant (seed, data split, epochs, hardware)
  4. Predict outcome — before running, state what you expect and why

Template for each ablation row:

| Run ID | Factor | Value | Fixed Config | Expected Outcome |
|--------|--------|-------|-------------|-----------------|

Step 3: Experiment Matrix

For multi-factor studies, use a structured matrix:

  1. Full factorial — if factors are few (≤3) and levels are few (≤3 each)
  2. Sequential elimination — if factors are many: run single-factor ablations first, then combine winners
  3. Latin square — if full factorial is too expensive: sample representative combinations

Always calculate total runs before committing:

Total runs = product of all factor levels
GPU hours = total runs × hours_per_run

Step 4: Resource Estimation

For each experiment plan, estimate:

  • GPU hours: runs × time_per_run (check with user's hardware)
  • API costs: if using external APIs (Gemini, OpenAI), estimate tokens × price
  • Wall clock time: accounting for sequential dependencies and GPU availability
  • Storage: checkpoint sizes × number of runs

Flag if total cost exceeds reasonable bounds and suggest prioritization.

Step 5: Config Stub Generation

Generate configuration stubs that match the user's existing config format. Read existing configs first to match:

  • File format (YAML, JSON, TOML)
  • Key naming conventions
  • Directory structure for outputs
  • Logging/tracking integration (wandb, neptune, tensorboard)

Step 6: Execution Plan

Create a concrete execution plan:

  1. Order runs by dependency (baselines first, then ablations)
  2. Identify which runs can be parallelized across GPUs
  3. Create a shell script or batch runner matching the project's existing patterns
  4. Include checkpointing strategy for long runs

Step 7: Analysis Plan

Before running, define how results will be analyzed:

  • Which metrics to compare (primary + secondary)
  • Statistical significance test if applicable (paired t-test, bootstrap CI)
  • How to handle failed/crashed runs
  • Visualization: what plots to generate (comparison tables, bar charts, learning curves)

Verification Checkpoints

Before finalizing the experiment plan:

  • [ ] Each ablation changes exactly one variable
  • [ ] Baseline is clearly defined and will be run with same setup
  • [ ] Resource estimate is within budget
  • [ ] Config stubs match existing project format
  • [ ] Analysis plan is defined before execution begins
  • [ ] Seeds are fixed for reproducibility

Output Format

Always produce:

  1. Experiment matrix table — all runs with their configurations
  2. Resource estimate — GPU hours, API costs, storage
  3. Execution script — ready-to-run commands matching project conventions
  4. Analysis plan — metrics, comparisons, visualizations

Related Skills

View on GitHub
GitHub Stars406
CategoryEducation
Updated18d ago
Forks34

Languages

Shell

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions