codex-ab
Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict
Install / Use
npx skills add oliver-kriska/claude-elixir-phoenix --skill codex-abInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
SecuritySupported Platforms
Our assessment of codex-ab
codex-ab scores 87/100 on our quality scale, 613th of 1,086 Security skills we index.
Its SKILL.md is 3.8 KB long, well organised into 11 sections with 4 code examples: a solid amount of guidance for an agent.
It has 560 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 2 days ago, so codex-ab is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
codex-ab compared with similar skills
All 4 of these similar skills score higher than codex-ab; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| codex-ab (this skill)by oliver-kriska | 87 | 560 | 2d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install codex-ab?
- Run
npx skills add oliver-kriska/claude-elixir-phoenix --skill codex-ab. The install tabs above show the steps for each supported agent. - Which AI agents does codex-ab work with?
- It is written for OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is codex-ab safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is codex-ab still maintained?
- The repository was last updated 2 days ago, so codex-ab is actively maintained.
Skill content
View source on GitHubname: codex-ab description: Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs. effort: medium argument-hint: "[base-branch]"
Codex Panel A/B (contributor instrument — verdict decided)
Answer one question with evidence: do dimension-focused codex passes find
real issues that one holistic codex exec review misses? Runs both on the
same diff, then classifies every focused finding against the holistic pass.
DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED. Fresh-only 1 real miss / 1 false positive plus one zero-value run at 4× cost — real misses did not outnumber FPs. Kept as contributor tooling (NOT distributed) for one possible retest: a UI-heavy diff with a single extra liveview-focused pass (2× cost). Scoreboard:
.claude/research/2026-07-03-codex-review-integration.md§7.
Usage
/codex-ab # A/B against main (~5 min, 4 codex runs)
/codex-ab develop # explicit base branch
Iron Laws
- FRESH DIFF ONLY — ask the user to confirm this branch has NOT been
codex-reviewed yet (cloud or
/phx:codex-loop). A drained diff returns NO FINDINGS everywhere and proves nothing — wasted quota - Verify every REAL MISS in the code before counting it — a focused finding only scores if the issue actually exists at that file:line
- Read ONLY the findings
.mdfiles — streams are diverted to.logfiles; never cat a log into context (10k+ lines each) - Exactly 4 codex runs, never re-run dimensions — bounded quota
- Persist the verdict — an unrecorded experiment is wasted quota
Workflow
Step 1: Preflight
Run command -v codex — missing → STOP with install hint. Then:
git status --shortdirty → warn (codex flags local dirt as findings)- Ask: "Has codex already reviewed this branch (PR review or codex-loop)?" If yes → STOP, explain the fresh-diff requirement (Iron Law 1)
Step 2: Run the A/B (background, ~5 min)
bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \
.claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M)
Use run_in_background — it runs 1 holistic codex exec review + 3
focused codex exec workers (security / ecto / liveview) in parallel,
all streams redirected. Do other work or wait; never poll.
Step 3: Classify
Read the 4 findings files (holistic.md, security.md, ecto.md,
liveview.md — small). For EACH focused finding:
| Class | Meaning | Test | |-------|---------|------| | DUPLICATE | Holistic already found it | Same file + same defect | | REAL MISS | Genuine issue holistic missed | Read the code at file:line — defect confirmed (Iron Law 2) | | FALSE POSITIVE | Manufactured, pre-existing, or wrong | Code check fails, or issue exists on base branch too |
Step 4: Verdict
Present:
## Codex Panel A/B — {branch} vs {base}
| dimension | findings | duplicate | real miss | false positive |
Holistic-only findings: {n}
Verdict this run: {REAL MISS count} real miss vs {FP count} false positive
Decision rule: build --codex-panel only if real misses outnumber false
positives across 2-3 fresh branches.
Write the verdict table to .claude/reviews/codex-ab-{date}/VERDICT.md.
Suggest repeating on the next 1–2 fresh branches before deciding.
Integration
fresh branch → /codex-ab (YOU ARE HERE) → verdict logged
├─ real misses win across runs → build /phx:review --codex-panel
└─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea
└─ OUTCOME 2026-07-10: this branch won — panel dropped
References
${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh— the 4-run harness- Related:
/phx:codex-loop(holistic fix loop),/phx:review --codex
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
