grill-jev
Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden problems before they ship.
Install / Use
npx skills add notque/vexjoy-agent --skill grill-jevInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of grill-jev
grill-jev scores 91/100 on our quality scale, 1023rd of 4,619 Development & Engineering skills we index (top 23%).
Its SKILL.md is 14 KB long, well organised into 19 sections with 5 code examples: a thorough specification that gives an agent plenty to work with.
It has 425 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated yesterday, so grill-jev is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
grill-jev compared with similar skills
All 4 of these similar skills score higher than grill-jev; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| grill-jev (this skill)by notque | 91 | 425 | 1d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.9k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install grill-jev?
- Run
npx skills add notque/vexjoy-agent --skill grill-jev. The install tabs above show the steps for each supported agent. - Which AI agents does grill-jev work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is grill-jev safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is grill-jev still maintained?
- The repository was last updated yesterday, so grill-jev is actively maintained.
Skill content
View source on GitHubname: grill-jev description: "Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden problems before they ship." user-invocable: true routing: force_route: true triggers: - grill - validate plan - validate spec - validate design - interrogate - stress test plan - stress test spec - 50 questions - deep validation - grill this - grill the plan - jev interrogation - sanity check plan - sanity check spec - check my plan - check my design - plan validation - spec validation - design validation - architecture review jev - verify plan - verify spec - challenge plan - challenge design - find holes - find gaps in plan not_for: "Writing new Jev programs (use building-with-jev). Routing requests (use do). Code review for quality/style (use review). Security scanning (use security). This skill runs a broad interrogation battery against a supplied artifact — it does not build Jev programs." pairs_with: - building-with-jev - review complexity: Complex category: meta allowed-tools:
- Read
- Write
- Edit
- Bash
- Glob
- Grep
Grill-Jev
Generate a context-specific Jev interrogation battery against a plan, spec, design, code artifact, or any topic. Questions are generated from the artifact itself — not from a static list — so they target the actual risks and gaps in what you provide.
When to invoke
- User says "grill this", "validate my plan", "find holes in my spec", "sanity check", "stress test this", "50 questions", "interrogate", "grill-jev"
- Automatically after any planning output. When a plan, spec, or design is produced — before execution begins — pass it through grill-jev. Any structured output with phases, steps, or checklist items qualifies.
- When the user says "does this look right", "approve this", "is this ready", "review my plan"
How it works
- Read the artifact — the plan, spec, design, or code being interrogated
- Generate questions — the LLM executing this skill produces up to 50 context-specific Jev questions tailored to the artifact. It writes the JSON battery to a temporary file; no secondary model call or API key is used. Questions use all structured shapes:
- Noul with
{true: {what, examples}, false: {what, examples}}criteria - Choice with
{what, not_for, examples}per option - Score with
{summary, signals}per level - Noul with array
compareinstructions when two state paths need comparison
- Noul with
- Send to Jev — evaluate the questions against the artifact as state. Use Vercel AI Gateway. Too much context is the most common failure: the artifact is resent with every batch, so split the battery by estimated tokens (
jev_limits.request_tokens), not by question count, keeping each request under the size target inskills/shared-patterns/jev-production-lessons.md. When the artifact alone passes the target, send the sections each question needs instead of the whole artifact.scripts/grill-jev.pydoes the splitting: it packs batches by tokens, and when a batch fails it halves and retries down to single questions, so one oversized or malformed question costs only its own answer. A question that still fails is listed as UNANSWERED (unknown, never a pass) and the run exits 2. If a tiny probe also fails, Jev is down and the run stops. The script rejects a malformed battery before sending (for example, Score levels must be a list). - Report findings — high-signal answers with suggested actions
Asking good questions
Question quality depends on context and framing. The rules below were validated through 3 iterative Jev loops.
Context inputs that improve questions:
| Input | Why it matters | Example | |---|---|---| | Audience | Who executes or approves this? SRE, junior dev, product owner, external team? | SREs need ops-specific questions; product owners need outcome and risk questions | | System context | What does the system actually do? What are its constraints? | Stateful vs. stateless changes need different risk questions | | Purpose | What decision does this artifact support? | Approval gate → binary questions; exploration → broader coverage | | Depth wanted | 15 sharp questions or 50 comprehensive ones? | Match count to stakes and complexity |
Pass context via --context "audience: SRE, system: stateful payment service, known constraint: cannot have >5min downtime".
Generation rules (Jev-validated):
-
Specific over generic — questions must name specific steps, systems, or claims in the artifact. "Does step 4's migration define a rollback safe to run under live traffic?" beats "does this have rollback?".
-
Mentions trigger deeper scrutiny, not shallower — when the artifact mentions a risk, gap, or uncertainty, generate MORE targeted questions about it. Acknowledgment is not mitigation. "This is risky" without a defined mitigation is itself a finding.
-
Audience weight — use audience context to focus questions. An SRE needs ops questions. A junior dev needs step-clarity questions. A product owner needs outcome questions.
-
Coverage balance — aim for breadth across relevant categories (completeness, feasibility, risk, scope, verification, consistency, reversibility, security, cost and throughput). Do not cluster all questions on one category. Skip a category only when the artifact has nothing that triggers it.
-
Scale by complexity — simple artifact (15-20 questions), medium (25-35), complex (40-50). Hard cap: 50.
Self-calibration loop — if findings feel generic or off-target, use Jev to improve:
- Run grill-jev — observe which findings feel shallow
- Ask Jev: "Which questions were not specific to this artifact? What context would have produced better questions?"
- Feed that context back via
--contextand re-run
Question categories
Generate questions covering these nine areas, weighted by what the artifact contains:
| Category | What it finds | |---|---| | Completeness | missing phases, undefined terms, unstated assumptions | | Feasibility | resource constraints, timeline, dependencies | | Risk & failure modes | what happens when each step fails | | Scope & boundaries | what's in/out, integration surfaces | | Verification | how do we know it worked, success criteria | | Consistency | internal contradictions, duplicate effort | | Reversibility | can we undo this, migration risk | | Security & safety | auth, data exposure, destructive operations | | Cost & throughput | calls, tokens, and requests per run and per second against the provider's documented rate limits; fan-out size; concurrency; retry policy; eval cost |
Plans that call Jev or another metered API
A plan can be complete, feasible, and safe and still fail in production because one run spends the provider's per-second limit. Grill it on arithmetic, not just on prose:
- Check it against the rules. Load
skills/shared-patterns/jev-production-lessons.mdand turn every unticked pre-ship checklist item into acost_Noul withreport_when: "false". Name the exact number in the question: "Does the plan keep every request at or under 4k tokens or a measured reliable size?", "Does it send each stage's requests at once, with an instance cap near floor(0.25 × 250,000 / tokens_per_request)?", "Does it set attempts, per-attempt timeout, run deadline, and a retry budget near 4 × requests × failure rate?" - Price it first. When the artifact calls Jev, build (or ask for) a JSON of every request one run sends and run
python3 scripts/jev-budget-check.py --payload run.json --concurrency C --concurrent-runs N --attempts A, plus--eval-cases Nfor any eval the plan runs. Afailis a high-signal finding on its own; awarngoes in the report. Put the check's summary in state underbudgetso battery questions can inspect it. - Ask about throughput explicitly. Include Cost & throughput questions (see
references/question-battery.md): does the plan state tokens per run and per second against the documented limits (250,000 input tokens per second and 1,200 requests per minute for Jev on 2026-09-22)? Does it fan full detail out over every unit, or cascade? Is in-flight concurrency capped? Do retries use jittered exponential backoff with a per-run budget? Is the eval priced and paced? Which errors mean "back off" on the production transport (Vercel AI Gateway reports upstream overload as 503)? - Treat "each request fits" as unproven. A plan that shows every request under the per-request limit has not shown the run fits. Look for the per-second number.
- Diagnoses need measurements. When the artifact explains a failure (an outage, a size cap, a bad payload), ask whether it measured the run's own rate and retry count and tested a small known-good request on the same transport before concluding.
Question generation prompt
Use this system prompt to generate the question battery:
You are generating a Jev question battery to interrogate an artifact.
The artifact is provided as state at key "artifact".
Generate between 20 and 50 questions. Scale the count to the artifact's
complexity — a 3-step bug fix needs fewer questions than a multi-service
migration plan.
For each question, choose the most appropriate Jev primitive:
- Noul: yes/no probability. Use when you want to know whether something
is true or absent. Always add structured criteria (true/false with
what + examples) when the boundary is non-obvious.
- Choice: pick one from a known set. Use for risk levels, categories,
reversibility classifications.
- Score: position on a spectrum. Use for completeness, timeline
realism, detection speed.
Use array compare instructions when two or more fields in the state
need to be compared side by side.
Focus questions on the actual content of the artifact. A plan with no
database steps needs no database migration questions. A plan with a
single deploy step needs no multi-service blast radius question.
Output a JSON dict of question_id -> question definition.
Running the battery
# File artifact
python3 scripts/grill-jev.py --file task_plan.md --mode plan
# Inline text
python3 scripts/grill-jev.py --text "$(cat task_plan.md)" --mode plan
# The executing LLM writes /tmp/grill-questions.json, then Jev evaluates it.
python3 scripts/grill-jev.py --file design.md --questions-file /tmp/grill-questions.json
The executing LLM owns question generation; the script only evaluates a supplied battery through Jev. Without --questions-file, it uses the static fallback battery. The compact guide in references/question-battery.md provides:
- Example shapes for the LLM question generator
- Fallback when generation is unavailable
- Test fixture for unit tests
Output
GRILL-JEV FINDINGS — mode: plan — 31 questions (generated)
============================================================
HIGH SIGNAL (requires attention):
[completeness/has_success_criteria] noul=0.89 TRUE — no measurable success criteria
→ add explicit success criteria: observable outcomes, not "it works"
[risk/migration_live_safety] noul=0.84 TRUE — migration runs against live traffic
→ add maintenance window or use online migration tool
CATEGORY SUMMARY:
completeness 2 findings risk 1 finding
OVERALL READINESS: 1.4/3 — partially complete; address findings before executing
Primitive shapes — quick reference
# Noul with structured criteria
"has_success_criteria": {
"type": "noul",
"instructions": {"question": "Does the plan define measurable success criteria?",
"inspect": "artifact"},
"criteria": {
"true": {"what": "Specific, observable outcomes are named",
"examples": ["all tests pass", "p95 latency < 200
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.9kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
