SkillAgentSearch skills...

ab-test-readout

Analyse a finished A/B test and write the readout — the result, whether it's statistically and practically significant, what it means, and the ship/no-ship call

Install / Use

npx skills add mohitagw15856/pm-claude-skills --skill ab-test-readout

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

82/100

Category

Marketing

Supported Platforms

Universal

Tags

Our assessment of ab-test-readout

ab-test-readout scores 82/100 on our quality scale, 422nd of 553 Marketing skills we index.

Its SKILL.md is 2.9 KB long, well organised into 11 sections and no code examples: a solid amount of guidance for an agent.

With 1,396 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
26/30
Structure
13/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 8 days ago, so ab-test-readout is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

ab-test-readout compared with similar skills

All 4 of these similar skills score higher than ab-test-readout; compare them before choosing.

SkillScoreStarsUpdatedFormat
ab-test-readout (this skill)by mohitagw15856821.4k8d agoSKILL.md
algorithmic-artby anthropics100177.9k10d agoSKILL.md
pptxby anthropics100177.9k10d agoSKILL.md
designby nextlevelbuilder100130.2k11d agoSKILL.md
ui-ux-pro-maxby nextlevelbuilder100130.2k11d agoSKILL.md

Frequently asked questions

How do I install ab-test-readout?
Run npx skills add mohitagw15856/pm-claude-skills --skill ab-test-readout. The install tabs above show the steps for each supported agent.
Which AI agents does ab-test-readout work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is ab-test-readout safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is ab-test-readout still maintained?
The repository was last updated 8 days ago, so ab-test-readout is actively maintained.

name: ab-test-readout description: "Analyse a finished A/B test and write the readout — the result, whether it's statistically and practically significant, what it means, and the ship/no-ship call. Use when asked to analyse experiment results, write an A/B test readout, interpret test data, or decide whether to ship a variant. Produces a clear verdict with the lift and confidence, segment cuts, the risks (peeking, novelty, sample), and a recommendation. Distinct from planning a test — this reads results."

A/B Test Readout Skill

The hard part of an experiment is the readout: not "B won" but "is this real, is it big enough to matter, and should we ship?" This skill turns results into an honest decision — and flags the ways A/B results lie.

Working from a brief

Given results (even partial), write the full readout anyway. If significance isn't provided, reason about it from the numbers and flag what's needed to confirm. Mark assumed figures. Never declare a winner without addressing significance and sample.

Required Inputs

Ask for (if not already provided):

  • The hypothesis and the primary metric
  • Results — control vs variant: conversions/rate, sample size per arm, duration
  • Guardrail metrics (revenue, retention, latency, complaints) that mustn't regress
  • Pre-registered decision rule (what would count as a win) if one exists

Output Format

1. Verdict (one line)

Ship / Don't ship / Inconclusive — keep running — with the headline number.

2. The result

| Metric | Control | Variant | Relative lift | Significant? | |---|---|---|---|---| | Primary | | | | p / CI | | Guardrail(s) | | | | |

State statistical significance (p-value / confidence interval) and practical significance (is the lift big enough to matter given the cost?).

3. Did it really win?

Address the ways A/B tests mislead:

  • Sample / power — was the test adequately powered, or under-sampled?
  • Peeking — was the call made early, inflating false positives?
  • Novelty / primacy — could the effect fade?
  • Segments — does the win hold across key segments, or is it driven by one?

4. Segment cuts

Where the effect is strong vs flat vs negative (new vs returning, platform, geography).

5. Recommendation & next step

Ship / iterate / re-run, plus what to monitor post-launch or what the follow-up test should isolate.

Quality Checks

  • [ ] Distinguishes statistical from practical significance
  • [ ] Checks guardrail metrics, not just the primary
  • [ ] Flags peeking, power, novelty, and segment-driven wins
  • [ ] Recommendation follows from the evidence, with a monitoring/next-test step
  • [ ] Doesn't declare a winner on an underpowered or peeked result

Anti-Patterns

  • "B won by 8%!" with no significance or sample size
  • Calling a result early (peeking) and shipping
  • Ignoring a guardrail regression because the primary went up
  • A statistically significant but practically meaningless lift treated as a win

Related Skills

View on GitHub
GitHub Stars1.4k
CategoryMarketing
Updated8d ago
Forks249

Languages

HTML

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions