exploratory-test
Execute multi-level exploratory testing of the app covering basic functionality, complex operations, adversarial testing, and cross-cutting scenarios, plus usability observations through a UX lens reported separately from defects. Deeper than /smoke-test
Install / Use
npx skills add tobihagemann/turbo --skill exploratory-testInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of exploratory-test
exploratory-test scores 88/100 on our quality scale, 1915th of 4,616 Development & Engineering skills we index (top 42%).
Its SKILL.md is 7.7 KB long, well organised into 19 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
It has 405 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 12 days ago, so exploratory-test is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-10-05. Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
exploratory-test compared with similar skills
All 4 of these similar skills score higher than exploratory-test; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| exploratory-test (this skill)by tobihagemann | 88 | 405 | 12d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.0k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.8k | 5d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install exploratory-test?
- Run
npx skills add tobihagemann/turbo --skill exploratory-test. The install tabs above show the steps for each supported agent. - Which AI agents does exploratory-test work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is exploratory-test safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is exploratory-test still maintained?
- The repository was last updated 12 days ago, so exploratory-test is actively maintained.
Skill content
View source on GitHubname: exploratory-test description: "Execute multi-level exploratory testing of the app covering basic functionality, complex operations, adversarial testing, and cross-cutting scenarios, plus usability observations through a UX lens reported separately from defects. Deeper than /smoke-test. Use when the user asks to "exploratory test", "test thoroughly", "test all scenarios", "deep test", "test edge cases", "test everything", "break it", "find bugs by testing", "test usability", or "check the UX while testing"."
Exploratory Test
Execute multi-level exploratory testing that goes beyond smoke testing to actively find bugs through escalating test scenarios.
Task Tracking
At the start, use TaskCreate to create a task for each step:
- Load or create test plan
- Determine testing approach
- Run
/user-experienceskill (when user-facing) - Run
/test-run-rulesskill - Execute tests by level
- Report
Step 1: Load or Create Test Plan
Resolve the test plan using these rules in order:
- Explicit path — If a file path was passed, use it
- Explicit slug — resolve to
.turbo/test-plans/<slug>.md - Anchoring artifact — If the work under test is anchored to a plan, resolve to
.turbo/test-plans/<that-slug>.mdwhen that file exists - Single file — Glob
.turbo/test-plans/*.md. If exactly one file exists, use it - Most recent — If multiple files exist, use the most recently modified
- Legacy fallback —
.turbo/test-plan.mdif.turbo/test-plans/does not exist - Nothing found — run the
/create-test-planskill first, then use the plan it writes
If multiple test plans exist and the most-recent choice is non-obvious, use AskUserQuestion to let the user pick from the candidates.
Read the resolved test plan and state its path.
Unless an explicit path or slug was passed, confirm the resolved plan still describes the work under test:
- Unavailable branch state — a scenario's steps require a branch that no longer resolves in the repository
- Completed prior run — every checkbox is already ticked and no recorded result is FAIL or PARTIAL
- Superseded context — the plan's Context section names work that changes merged since the plan was written have reversed or removed
When a signal fires, output the signal and the scenarios it affects as text. For a superseded Context, name the scenarios that exercise the reversed or removed work. Then use AskUserQuestion to offer:
- Regenerate — run the
/create-test-planskill with the resolved path, and use the plan it writes - Execute anyway — the signal is a false positive
- Pick another plan — resolve to a different test plan file, then confirm that plan against these same signals
If the user specifies a narrower scope, filter the plan to relevant scenarios rather than executing all of them. Reserve filtering for that case: a superseded plan keeps scenarios that each look plausible alone, so trimming it preserves the wrong ones.
Step 2: Determine Testing Approach
Use the approach specified in the test plan. If the plan does not specify one, determine it using the same logic as /create-test-plan Step 2.
Step 3: Run /user-experience Skill (When User-Facing)
If the app has a user-facing surface (UI, screens, commands, messages, or any behavior a user sees or does), run the /user-experience skill to load the UX lens before executing tests, so usability concerns surface while interacting with the app. When it is unclear whether the surface is user-facing, use AskUserQuestion to ask rather than skipping silently. Skip this step for test targets with no user-facing behavior (internal library or infrastructure).
Step 4: Run /test-run-rules Skill
Run the /test-run-rules skill to load the rules for launching and driving the app.
Step 5: Execute Tests by Level
Work through each level sequentially. Complete all tests in a level before moving to the next.
Execution Loop (Per Test)
- Set up the preconditions described in the test scenario
- Perform the exact steps
- Capture the result (screenshot, output, or state observation)
- Compare against the expected outcome
- Record PASS, FAIL, or PARTIAL with details
- When the UX lens is loaded, note any usability observation it surfaces, kept separate from the verdict
Record a scenario the test run rules leave blocked or inconclusive as PARTIAL, naming what is unproven and why.
When the scenario's output is consumed by another system, withhold PASS until that system accepts it. Decoding a token, reading a response body, or confirming a row exists shows only that the artifact was produced. Stand up the consumer under the same isolation and cleanup rules as any other service this run starts, and exercise its own flow. When standing it up is not possible, record PARTIAL and name which half is unproven. PARTIAL counts as not passed everywhere a verdict is tallied or gated.
Level Progression
- Level 1: Basic Functionality — If any Level 1 test does not pass, report early and use
AskUserQuestionto ask whether to continue. Basic failures may indicate the feature is too broken for deeper testing. - Level 2: Complex Operations — Execute all tests regardless of individual failures.
- Level 3: Adversarial Testing — Execute all tests. Failures here are expected and valuable.
- Level 4: Cross-Cutting Scenarios — Execute all tests.
If a project-specific testing skill or MCP tool was identified in Step 2, use that. The paths below are fallbacks.
Web App Path
Start or reuse a dev server under the test run rules. If /agent-browser is available, run the /agent-browser skill. Otherwise, use claude-in-chrome MCP to interact with the app.
UI/Native App Path
Launch the app. Use computer-use MCP to interact with the UI.
CLI Path
Run commands directly.
Step 6: Report
Present results organized by level:
Exploratory Test Results:
## Level 1: Basic Functionality (X/Y passed)
- [PASS] Test name: description — [substitution, when one was driven]
- [FAIL] Test name: description — [what went wrong]
- [PARTIAL] Test name: description — [what is unproven, and why]
## Level 2: Complex Operations (X/Y passed)
- [PASS] Test name: description — [substitution, when one was driven]
- [FAIL] Test name: description — [what went wrong]
- [PARTIAL] Test name: description — [what is unproven, and why]
## Level 3: Adversarial Testing (X/Y passed)
- [PASS] Test name: description — [substitution, when one was driven]
- [FAIL] Test name: description — [what went wrong]
- [PARTIAL] Test name: description — [what is unproven, and why]
## Level 4: Cross-Cutting Scenarios (X/Y passed)
- [PASS] Test name: description — [substitution, when one was driven]
- [FAIL] Test name: description — [what went wrong]
- [PARTIAL] Test name: description — [what is unproven, and why]
Overall: X/Y passed across all levels
Report usability observations from the UX lens below the level results, separately from the defects. A scenario can pass every functional check and still surface a usability concern.
## Usability Observations
- [UX] <observation> — names the UX context it touches (Understanding, Bridging, or Flowing) and the goal mismatch or friction it creates
For each failure, include the relevant screenshot, output, or state observation.
When the change under test spans several repositories, add a per-repo view of the findings below the usability observations, naming a suggested fix site for each.
Update the resolved test plan file by checking off completed tests and annotating results.
Then use the TaskList tool and proceed to any remaining task.
Rules
- To diagnose failures, run the
/investigateskill on the test report.
Related Skills
ai-job-search
45.0kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.8kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
