pdf-processor
Perform basic local PDF operations (merge, split, extract pages/text/tables, create) when users request offline PDF processing without external services.
Install / Use
npx skills add aipoch/medical-research-skills --skill pdf-processorInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of pdf-processor
pdf-processor scores 92/100 on our quality scale, 248th of 881 Content & Media skills we index (top 29%).
Its SKILL.md is 6.6 KB long, well organised into 19 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
With 1,916 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 12 days ago, so pdf-processor is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-30. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
pdf-processor compared with similar skills
All 4 of these similar skills score higher than pdf-processor; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pdf-processor (this skill)by aipoch | 92 | 1.9k | 12d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.2k | 14d ago | CLAUDE.md |
| siyuanby siyuan-note | 100 | 46.6k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 7d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 7d ago | SKILL.md |
Frequently asked questions
- How do I install pdf-processor?
- Run
npx skills add aipoch/medical-research-skills --skill pdf-processor. The install tabs above show the steps for each supported agent. - Which AI agents does pdf-processor work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is pdf-processor safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pdf-processor still maintained?
- The repository was last updated 12 days ago, so pdf-processor is actively maintained.
Skill content
View source on GitHubname: pdf-processor description: Perform basic local PDF operations (merge, split, extract pages/text/tables, create) when users request offline PDF processing without external services. license: MIT author: AIPOCH
When to Use
- You need to merge multiple PDFs into a single document offline.
- You need to split a PDF into per-page files for review, annotation, or distribution.
- You need to extract a specific page range (e.g.,
1-3,5) into a new PDF. - You need to extract selectable text from a PDF into a
.txtfile without using any cloud API. - You need to extract tables from PDFs into CSV files locally.
Key Features
- Merge PDFs: Combine multiple input PDFs into one output PDF.
- Split PDFs: Export each page as an individual PDF.
- Extract Pages: Create a new PDF from selected pages using a page-range expression.
- Extract Text: Export extracted text to a plain text file.
- Extract Tables: Export detected tables to CSV files (may produce empty CSVs on pages without tables).
- Create PDF: Generate a simple PDF from a text file.
- Local-only execution: No network access, no external APIs, no credentials.
Dependencies
- Python: 3.9+
Install Python dependencies:
pip install -r scripts/requirements.txt
Example Usage
More examples may be available in
references/examples.md.
Merge PDFs
python scripts/pdf_tool.py \
--operation merge \
--inputs "a.pdf" "b.pdf" \
--output "out.pdf"
Split a PDF into single-page PDFs
python scripts/pdf_tool.py \
--operation split \
--inputs "input.pdf" \
--output "out_dir"
Extract specific pages into a new PDF
python scripts/pdf_tool.py \
--operation extract-pages \
--inputs "input.pdf" \
--pages "1-3,5" \
--output "extracted.pdf"
Extract text to a .txt file
python scripts/pdf_tool.py \
--operation extract-text \
--inputs "input.pdf" \
--output "output.txt"
Extract tables to CSV files
python scripts/pdf_tool.py \
--operation extract-tables \
--inputs "input.pdf" \
--output "tables_out_dir"
Create a PDF from a text file
python scripts/pdf_tool.py \
--operation create \
--inputs "input.txt" \
--output "created.pdf"
Implementation Details
- Execution model: Local CLI tool/script (
scripts/pdf_tool.py) that reads from provided input paths and writes only to the specified output path. - Supported operations:
merge: Concatenates PDFs in the order provided via--inputs.split: Writes one PDF per page (output is typically a directory path).extract-pages: Uses--pagesto select pages and writes a new PDF.extract-text: Extracts selectable text; pages with no extractable text may yield empty lines.extract-tables: Attempts table detection/extraction; pages without tables may produce empty CSV outputs.create: Produces a simple PDF from a text input.
- Page range format:
--pages "1-3,5"where page numbering starts at 1. - Failure handling:
- Invalid or out-of-range pages in
--pagesare ignored. - Text extraction preserves structure minimally; no OCR is performed.
- Table extraction quality depends on the PDF's structure; complex layouts may not reconstruct well.
- Invalid or out-of-range pages in
- Security constraints:
- No network calls.
- No external services.
- Reads only specified input files and writes only to the specified output location.
- Success criteria:
- Output file(s) exist at the specified location.
- Page counts and ordering match the requested operation.
- No writes occur outside the specified output path.
When Not to Use
- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
Required Inputs
- A clearly specified task goal aligned with the documented scope.
- All required files, identifiers, parameters, or environment variables before execution.
- Any domain constraints, formatting requirements, and expected output destination if applicable.
Recommended Workflow
- Validate the request against the skill boundary and confirm all required inputs are present.
- Select the documented execution path and prefer the simplest supported command or procedure.
- Produce the expected output using the documented file format, schema, or narrative structure.
- Run a final validation pass for completeness, consistency, and safety before returning the result.
Output Contract
- Return a structured deliverable that is directly usable without reformatting.
- If a file is produced, prefer a deterministic output name such as
pdf_processor_result.mdunless the skill documentation defines a better convention. - Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
Validation and Safety Rules
- Validate required inputs before execution and stop early when mandatory fields or files are missing.
- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
- Keep the output safe, reproducible, and within the documented scope at all times.
Failure Handling
- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
- If partial output is returned, label it clearly and identify which checks could not be completed.
Input Validation
This skill accepts requests that match the documented purpose of pdf-processor and include enough context to complete the workflow safely.
Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:
pdf-processoronly handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.
Quick Validation
Run this minimal verification path before full execution when possible:
python scripts/pdf_tool.py --help
Expected output format:
Result file: pdf_processor_result.md
Validation summary: PASS/FAIL with brief notes
Assumptions: explicit list if any
Related Skills
Agent-Reach
86.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
siyuan
46.6kAn open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
