ab-test-analyzer
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation
Install / Use
npx skills add irinabuht12-oss/marketing-skills --skill ab-test-analyzerInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of ab-test-analyzer
ab-test-analyzer scores 89/100 on our quality scale, 1560th of 4,615 Development & Engineering skills we index (top 34%).
Its SKILL.md is 5.4 KB long, well organised into 29 sections with 4 code examples: a solid amount of guidance for an agent.
With 1,836 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 12 days ago, so ab-test-analyzer is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-06. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
ab-test-analyzer compared with similar skills
All 4 of these similar skills score higher than ab-test-analyzer; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| ab-test-analyzer (this skill)by irinabuht12-oss | 89 | 1.8k | 12d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.1k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.8k | 5d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install ab-test-analyzer?
- Run
npx skills add irinabuht12-oss/marketing-skills --skill ab-test-analyzer. The install tabs above show the steps for each supported agent. - Which AI agents does ab-test-analyzer work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is ab-test-analyzer safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is ab-test-analyzer still maintained?
- The repository was last updated 12 days ago, so ab-test-analyzer is actively maintained.
Skill content
View source on GitHubname: ab-test-analyzer description: Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation. Use when feeding test results, checking statistical significance, calculating sample sizes, analyzing experiment outcomes, or generating next test ideas based on results. Platform: Google and Meta. metadata: platform: Google and Meta
A/B Test Analyzer
Evaluate A/B test results with statistical rigor and generate actionable next steps.
Process
- Collect test data - Variations, sample sizes, conversions, time period
- Check validity - Runtime, sample size, peeking issues
- Calculate significance - Z-score, p-value, confidence interval
- Segment analysis - Device, source, new vs returning
- Interpret results - Statistical vs practical significance
- Generate hypotheses - "Why it worked" and next test ideas
Sample Size Calculator
n = 2 × (Zα/2 + Zβ)² × p(1-p) / δ²
Where:
- n = sample size per variation
- Zα/2 = 1.96 (95% confidence) or 2.58 (99%)
- Zβ = 0.84 (80% power) or 1.28 (90%)
- p = baseline conversion rate
- δ = minimum detectable effect (absolute)
Quick Reference (95% confidence, 80% power): | Baseline CR | 10% Relative MDE | 20% Relative MDE | |-------------|------------------|------------------| | 2% | 78,000/var | 19,500/var | | 5% | 30,000/var | 7,500/var | | 10% | 14,300/var | 3,600/var |
Significance Calculation
Z-score = (pB - pA) / √(SE²_A + SE²_B)
SE = √(p × (1-p) / n)
If |Z| > 1.96 → Significant at 95%
If |Z| > 2.58 → Significant at 99%
Validity Checklist
- [ ] Ran for minimum 14 days (day-of-week effects)
- [ ] Met required sample size before concluding
- [ ] No "peeking" at results during test (inflates false positives to 40%+)
- [ ] Traffic split was random and consistent
- [ ] No external factors (holidays, outages, campaigns)
Bayesian vs Frequentist
| Use Case | Approach | |----------|----------| | Large samples (>10K), yes/no decision | Frequentist | | Continuous monitoring, early stopping | Bayesian | | Regulatory/legal requirements | Frequentist | | Want probability statements | Bayesian |
Output Format
## A/B Test Analysis
### Test Summary
- **Test Name**: [Name]
- **Hypothesis**: [What you expected]
- **Runtime**: [X days]
- **Sample Size**: Control [n] / Variant [n]
### Results
| Metric | Control | Variant | Lift | Significant? |
|--------|---------|---------|------|--------------|
| Conversion Rate | X% | Y% | +Z% | ✅/❌ |
### Statistical Details
- Z-score: [X]
- P-value: [X]
- Confidence Interval: [X% to Y%]
- Power achieved: [X]%
### Validity Assessment
- ✅/❌ Minimum runtime met
- ✅/❌ Sample size requirement met
- ✅/❌ No external confounds identified
### Segment Breakdown
| Segment | Control CR | Variant CR | Lift | Notes |
|---------|-----------|------------|------|-------|
| Mobile | | | | |
| Desktop | | | | |
### Interpretation
[Statistical significance vs practical significance discussion]
### "Why It Worked" Hypotheses
1. [Hypothesis about mechanism]
2. [Alternative explanation]
### Next Test Ideas
1. [Follow-up test based on learnings]
2. [Adjacent hypothesis to explore]
### Confidence Level: [HIGH/MEDIUM/LOW]
Example
Input: "Tested new headline. Control: 10,000 visitors, 250 conversions. Variant: 10,000 visitors, 300 conversions. Ran for 7 days."
Output:
## A/B Test Analysis
### Test Summary
- **Hypothesis**: New headline increases conversions
- **Runtime**: 7 days ⚠️
- **Sample Size**: Control 10,000 / Variant 10,000
### Results
| Metric | Control | Variant | Lift | Significant? |
|--------|---------|---------|------|--------------|
| Conversion Rate | 2.5% | 3.0% | +20% | ✅ Yes (95%) |
### Statistical Details
- Z-score: 2.28
- P-value: 0.023
- Confidence Interval: +2.3% to +37.7%
- Power achieved: 62% ⚠️
### Validity Assessment
- ❌ Minimum runtime NOT met (7 days < 14 days recommended)
- ⚠️ Sample size marginal for 20% MDE
- ❓ Cannot assess external confounds without more context
### Interpretation
Result is **statistically significant** but validity concerns exist:
1. 7-day runtime may miss day-of-week patterns
2. Wide confidence interval (+2% to +38%) indicates uncertainty
3. Recommend extending test 7 more days to confirm
### "Why It Worked" Hypotheses
1. New headline more clearly communicates value proposition
2. Specificity/numbers in headline increased credibility
### Next Test Ideas
1. Test headline variations that emphasize the winning element
2. Apply same messaging pattern to subheadline
### Confidence Level: MEDIUM
Statistical significance achieved, but short runtime reduces confidence.
Guidelines
- Never declare a winner without checking validity
- Distinguish statistical significance from practical significance
- If test ran <7 days, always recommend extending
- If sample size insufficient, calculate required runtime to reach it
- Ask for segment data if not provided - results often differ by device/source
Data access (Ryze MCP)
This skill works best with live account data. Connect the free Ryze MCP once and Claude reads your Google Ads, Meta Ads, GA4 and Search Console directly:
- claude.ai / Claude Desktop: Settings → Connectors → Add custom connector →
https://connector.get-ryze.ai/mcp - Claude Code:
claude mcp add ryze --transport http https://connector.get-ryze.ai/mcp - Cursor: Settings → MCP → add the same URL
Setup guide: https://www.get-ryze.ai/how-to-connect-claude-to-google-meta-ads-mcp
Related Skills
ai-job-search
45.1kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.8kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
