SkillAgentSearch skills...

academic-pdf-redaction

Redact text from PDF documents for blind review anonymization

Install / Use

npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

85/100

Supported Platforms

Universal

Our assessment of academic-pdf-redaction

academic-pdf-redaction scores 85/100 on our quality scale, 713th of 1,209 Content & Media skills we index.

Its SKILL.md is 3.6 KB long, well organised into 9 sections with 3 code examples: a solid amount of guidance for an agent.

With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
26/30
Structure
18/20
Description
12/15
Adoption
14/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated about 2 months ago, so academic-pdf-redaction is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

academic-pdf-redaction compared with similar skills

All 4 of these similar skills score higher than academic-pdf-redaction; compare them before choosing.

SkillScoreStarsUpdatedFormat
academic-pdf-redaction (this skill)by benchflow-ai851.8k2mo agoSKILL.md
Agent-Reachby Panniantong10090.5k19d agoCLAUDE.md
headroomby headroomlabs-ai10074.4k1d agoCLAUDE.md
Scraplingby D4Vinci10085.6ktodayMCP Server
crawl4aiby unclecode10084.8k9d agoMCP Server

Frequently asked questions

How do I install academic-pdf-redaction?
Run npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction. The install tabs above show the steps for each supported agent.
Which AI agents does academic-pdf-redaction work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is academic-pdf-redaction safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is academic-pdf-redaction still maintained?
The repository was last updated about 2 months ago, so academic-pdf-redaction is actively maintained.

name: academic-pdf-redaction description: Redact text from PDF documents for blind review anonymization

PDF Redaction for Blind Review

Redact identifying information from academic papers for blind review.

CRITICAL RULES

  1. PRESERVE References section - Self-citations MUST remain intact
  2. ONLY redact specific text matches - Never redact entire pages/regions
  3. VERIFY output - Check that 80%+ of original text remains

Common Pitfalls to AVOID

# ❌ WRONG - This removes ALL text from the page:
for block in page.get_text("blocks"):
    page.add_redact_annot(fitz.Rect(block[:4]))

# ❌ WRONG - Drawing rectangles over text:
page.draw_rect(fitz.Rect(0, 0, 600, 100), fill=(0,0,0))

# ✅ CORRECT - Only redact specific search matches:
for rect in page.search_for("John Smith"):
    page.add_redact_annot(rect)

Patterns to Redact (Before References Only)

IMPORTANT: Use FULL names/phrases, not partial matches!

  • ✅ "John Smith" (full name)
  • ❌ "Smith" (partial - would incorrectly match "Smith et al." citations in References)
  1. Author names - FULL names only (e.g., "John Smith", not just "Smith")
  2. Affiliations - Universities, companies (e.g., "Duke University")
  3. Email addresses - Pattern: *@*.edu, *@*.com
  4. Venue names - Conference/workshop names (e.g., "ICML 2024", "ICML Workshop")
  5. arXiv identifiers - Pattern: arXiv:XXXX.XXXXX
  6. DOIs - Pattern: 10.XXXX/...
  7. Acknowledgement names - Names in "Acknowledgements" section
  8. Equal contribution footnotes - e.g., "Equal contribution", "* Equal contribution"

PyMuPDF (fitz) - Recommended Approach

import fitz
import os

def redact_with_pymupdf(input_path: str, output_path: str, patterns: list[str]):
    """Redact specific patterns from PDF using PyMuPDF."""
    doc = fitz.open(input_path)
    original_len = sum(len(p.get_text()) for p in doc)

    # Find References page - stop redacting there
    references_page = None
    for i, page in enumerate(doc):
        if "references" in page.get_text().lower():
            references_page = i
            break

    for page_num, page in enumerate(doc):
        if references_page is not None and page_num >= references_page:
            continue  # Skip References section

        for pattern in patterns:
            # ONLY redact exact search matches
            for rect in page.search_for(pattern):
                page.add_redact_annot(rect, fill=(0, 0, 0))
        page.apply_redactions()

    os.makedirs(os.path.dirname(output_path), exist_ok=True)
    doc.save(output_path)
    doc.close()

    # MUST verify after saving
    verify_redaction(input_path, output_path)

REQUIRED: Verification Function

Always run this after ANY redaction to catch errors early:

import fitz

def verify_redaction(original_path, output_path):
    """Verify redaction didn't corrupt the PDF."""
    orig = fitz.open(original_path)
    redc = fitz.open(output_path)

    orig_len = sum(len(p.get_text()) for p in orig)
    redc_len = sum(len(p.get_text()) for p in redc)

    print(f"Original: {len(orig)} pages, {orig_len} chars")
    print(f"Redacted: {len(redc)} pages, {redc_len} chars")
    print(f"Retained: {redc_len/orig_len:.1%}")

    # DEFENSIVE CHECKS - fail fast if something went wrong
    if len(redc) != len(orig):
        raise ValueError(f"Page count changed: {len(orig)} -> {len(redc)}")
    if redc_len < 1000:
        raise ValueError(f"PDF corrupted: only {redc_len} chars remain!")
    if redc_len < orig_len * 0.7:
        raise ValueError(f"Too much removed: kept only {redc_len/orig_len:.0%}")

    orig.close()
    redc.close()
    print("✓ Verification passed")

Related Skills

View on GitHub
GitHub Stars1.8k
CategoryContent
Updated2mo ago
Forks368

Languages

PDDL

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions