eval-suite
Batch-evaluate a whole set of content pieces — files, a directory, or pasted blocks — in one run, producing a ranked portfolio quality report: grade distribution, per-dimension averages, systemic issues, a prioritized revision list, and auto-rejects below threshold.
Install / Use
npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suiteInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
MarketingSupported Platforms
Our assessment of eval-suite
eval-suite scores 82/100 on our quality scale, 521st of 610 Marketing skills we index.
Its SKILL.md is 8.9 KB long, split into 6 sections and no code examples: a thorough specification that gives an agent plenty to work with.
It has 832 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 26 days ago, so eval-suite is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
eval-suite compared with similar skills
All 4 of these similar skills score higher than eval-suite; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| eval-suite (this skill)by indranilbanerjee | 82 | 832 | 26d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.1k | 18d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install eval-suite?
- Run
npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suite. The install tabs above show the steps for each supported agent. - Which AI agents does eval-suite work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is eval-suite safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is eval-suite still maintained?
- The repository was last updated 26 days ago, so eval-suite is actively maintained.
Skill content
View source on GitHubname: eval-suite description: "Batch-evaluate a whole set of content pieces — files, a directory, or pasted blocks — in one run, producing a ranked portfolio quality report: grade distribution, per-dimension averages, systemic issues, a prioritized revision list, and auto-rejects below threshold. Triggers on "/digital-marketing-pro:eval-suite", "score our whole content library", "quality-check all campaign assets before launch", "evaluate these 5 drafts together", "which deliverables are weakest". Runs eval-runner.py per item, logs every score to the quality tracker for trend analysis, and reads the brand profile and guidelines for scoring context."
/digital-marketing-pro:eval-suite
Purpose
Batch evaluation across multiple content pieces to produce a portfolio-level quality assessment. Evaluate an entire content library, all assets in a campaign, or a set of deliverables in one run. Instead of evaluating content one piece at a time, this command processes everything together and delivers a holistic view of content quality.
The output includes content rankings, per-dimension analysis, overall quality distribution, common issues across the set, and a prioritized revision list. This is the command to use before a campaign launch (to catch weak assets before they go live), during a content audit (to assess library health), or after a production sprint (to quality-check all deliverables at once). Every evaluation is logged to the quality tracker for longitudinal trend analysis.
Input Required
The user must provide (or will be prompted for):
- Content sources: One or more of the following:
- A list of file paths (e.g., "evaluate these 5 files: email-v1.txt, email-v2.txt, landing-page.html, ad-copy-fb.txt, ad-copy-google.txt")
- A directory path (e.g., "evaluate everything in /campaign-q1-assets/") — all text-based files in the directory will be included
- Multiple inline content blocks with labels (e.g., "Evaluate these: [Label: Homepage Hero] content... [Label: Email Subject] content...")
- Content type: Optional — applied globally (e.g., "these are all email subject lines") or specified per item. If omitted, the evaluator will infer type from content characteristics
- Evidence file: Optional — shared context document (brief, strategy doc, audience research) applied across all evaluations for more relevant scoring
- Evaluation depth: Optional —
quick(default, faster per-item evaluation) orfull(comprehensive evaluation with detailed per-dimension commentary per item). Quick is recommended for sets larger than 10 items; full for critical campaign assets - Auto-reject threshold: Optional — composite score below which content is flagged as needing mandatory revision (default: 60)
- Comparison baseline: Optional — a previous eval-suite run ID to compare against, showing improvement or regression per piece
Process
- Load brand context: Read
~/.claude-marketing/brands/_active-brand.jsonfor the active slug, then load~/.claude-marketing/brands/{slug}/profile.json. Apply brand voice, compliance rules for target markets (skills/context-engine/compliance-rules.md), and industry context. Check for guidelines at~/.claude-marketing/brands/{slug}/guidelines/_manifest.json— if present, load restrictions and relevant category files (voice-and-tone, messaging, channel styles). Check for custom templates at~/.claude-marketing/brands/{slug}/templates/. Check for agency SOPs at~/.claude-marketing/sops/. If no brand exists, ask: "Set up a brand first (/digital-marketing-pro:brand-setup)?" — or proceed with defaults. - Enumerate all content items: Resolve the provided sources into a flat list of content items. For directory paths, scan for text-based files (.txt, .md, .html, .csv rows). For inline content, parse labels and content blocks. Assign a label to each item (filename, provided label, or auto-generated index). Report the total item count to the user before proceeding and confirm if the set is larger than 25 items (to set expectations on processing time).
- Evaluate each content item: For each item in the set, run
python "${CLAUDE_PLUGIN_ROOT}/scripts/eval-runner.py" --brand {slug} --action run-quick --file "{path}" --content-type "{type}"for file items (use--text "{content}"instead of--filefor inline content blocks; use--action run-fullif the user requested comprehensive depth). Pass--evidence "{evidence_path}"if an evidence file was provided. Collect the per-dimension scores (content_quality, brand_voice, hallucination_risk, claim_verification, output_structure, readability) and composite score for each item. - Log each evaluation: For every evaluated item, run
python "${CLAUDE_PLUGIN_ROOT}/scripts/quality-tracker.py" --brand {slug} --action log-eval --content-type "{type}" --data '{"label": "{label}", "scores": {scores_json}, "suite_id": "{suite_run_id}"}'to persist results for longitudinal tracking. The suite-id groups all items from this batch together. - Aggregate results: Compute portfolio-level statistics:
- Average composite score across all items
- Score distribution — count of items in each grade band (90+: Excellent, 80-89: Strong, 70-79: Good, 60-69: Needs Work, <60: Auto-reject)
- Per-dimension portfolio averages — identify which quality dimensions are consistently strong or weak across the entire set
- Standard deviation to assess consistency (high deviation means uneven quality)
- Rank all content pieces: Sort items from highest to lowest composite score. Present the full ranked list with scores, grades, and content type labels.
- Identify common issues: Analyze the per-dimension scores across all items to find patterns — e.g., "7 of 12 items score below 70 on claim_verification" or "hallucination_risk scores are consistently 15+ points below content_quality scores." These systemic patterns indicate process or template issues rather than individual content problems.
- Generate prioritized revision list: Sort items that need revision by potential impact. Prioritize items that are (a) below the auto-reject threshold, (b) high-visibility content types (landing pages, ads) with below-average scores, or (c) items where a single dimension drags down an otherwise strong composite. For each item on the revision list, specify which dimension(s) to focus on and what kind of improvement is needed.
- Compare against baseline (if provided): If the user provided a previous suite run ID, retrieve both the current and baseline suite scores from the quality tracker using
python "${CLAUDE_PLUGIN_ROOT}/scripts/quality-tracker.py" --brand {slug} --action get-summaryfor each suite period. Then compute per-item and portfolio-level deltas yourself by matching items across the two runs by label/content-type and calculating score differences. Present results as improved, regressed, or unchanged per item and overall.
Output
A structured portfolio quality assessment containing:
- Portfolio summary: Total piece count, average composite score, grade distribution (Excellent/Strong/Good/Needs Work/Auto-reject counts), overall portfolio grade, consistency score (based on standard deviation)
- Ranked content list: All items sorted best to worst — each with label, content type, composite score, grade, and a one-line quality summary
- Top performers: The 3 highest-scoring items with specific notes on what makes them strong — useful as internal benchmarks or templates
- Per-dimension portfolio analysis: Average score per dimension across the full set (content_quality, brand_voice, hallucination_risk, claim_verification, output_structure, readability), identifying the strongest and weakest dimensions with specific observations (e.g., "brand_voice averages 88 across the set — voice guidelines are being followed well. claim_verification averages 62 — sources and supporting evidence are frequently missing.")
- Common issues report: Systemic patterns found across multiple items — these indicate process-level problems worth fixing at the template or brief stage rather than per-item revision
- Prioritized revision list: Items most in need of revision, sorted by impact, with specific guidance on which dimensions to improve and what kind of changes are needed
- Auto-reject list: Items scoring below the threshold with specific reasons and mandatory revision flags
- Baseline comparison (if applicable): Per-item deltas and portfolio-level improvement/regression metrics
- Recommendations: Actionable next steps — which items to revise first, which process improvements would lift the entire portfolio, and whether any content types consistently underperform (suggesting brief or template issues)
Agents Used
- quality-assurance -- Evaluates each content piece across all quality dimensions, maintains scoring consistency across the batch, identifies systemic quality patterns, generates portfolio-level insights, and produces the prioritized revision recommendations
Related Skills
Agent-Reach
90.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
