pdf-extract
Extract PDF selectable text and full-page or segmented page images (including tables) into Markdown with per-page headings and image links; use when you need both readable text and page visuals for PPT creation, review, or analysis.
Install / Use
npx skills add aipoch/medical-research-skills --skill pdf-extractInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of pdf-extract
pdf-extract scores 92/100 on our quality scale, 247th of 881 Content & Media skills we index (top 29%).
Its SKILL.md is 7.1 KB long, well organised into 15 sections with 4 code examples: a thorough specification that gives an agent plenty to work with.
With 1,916 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 12 days ago, so pdf-extract is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-30. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
pdf-extract compared with similar skills
All 4 of these similar skills score higher than pdf-extract; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pdf-extract (this skill)by aipoch | 92 | 1.9k | 12d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.2k | 14d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.6k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.5k | 4d ago | MCP Server |
Frequently asked questions
- How do I install pdf-extract?
- Run
npx skills add aipoch/medical-research-skills --skill pdf-extract. The install tabs above show the steps for each supported agent. - Which AI agents does pdf-extract work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is pdf-extract safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pdf-extract still maintained?
- The repository was last updated 12 days ago, so pdf-extract is actively maintained.
Skill content
View source on GitHubname: pdf-extract description: Extract PDF selectable text and full-page or segmented page images (including tables) into Markdown with per-page headings and image links; use when you need both readable text and page visuals for PPT creation, review, or analysis. license: MIT author: AIPOCH
Validation Shortcut
Run this minimal command first to verify the supported execution path:
python scripts/extract_pdf.py --help
When to Use
- Converting a PDF report/paper into Markdown while preserving page structure (
## Page XX) for easy navigation. - Preparing PPT or design materials where you need both extracted text and page screenshots/blocks (tables, figures, diagrams).
- Reviewing scanned or mixed PDFs and filtering out "text-only" screenshots to keep only meaningful visuals.
- Building a dataset for downstream analysis where each page's text and images must be linked and traceable.
- Creating a lightweight "PDF-to-Markdown" archive with per-page headings and image references.
Key Features
- Extracts selectable PDF text and normalizes paragraphs (collapses line breaks into readable paragraphs).
- Writes Markdown with a document title and per-page sections (
# <filename>,## Page XX). - Supports multiple image extraction modes:
- segment (default): renders segmented page blocks (useful for tables/figures).
- embedded: extracts embedded images from the PDF.
- page: renders full pages as images.
- Optional post-filters to reduce noisy images:
- Filter text-heavy images via OCR (
--filter-text). - Drop images that match extracted page text (likely screenshots of text) (
--filter-match). - Drop images overlapping PDF text blocks without OCR (
--filter-pdf-text).
- Filter text-heavy images via OCR (
- Produces stable, per-page image links in Markdown for easy referencing.
Dependencies
pdfplumber(version not specified)pymupdf(version not specified)pytesseract(version not specified; required only when--filter-text onor--filter-match on)
Example Usage
python scripts/extract_pdf.py \
--input input.pdf \
--output output.md \
--image-dir images \
--image-mode segment \
--filter-text on \
--text-threshold 0.25 \
--text-lang eng \
--filter-match on \
--match-lang eng \
--match-min-len 30 \
--filter-pdf-text on \
--pdf-text-threshold 0.1
Expected Markdown structure:
- Document title:
# <filename> - Per-page section:
## Page XX - Text paragraphs (normalized)
- Image links, depending on mode:
- Segmented blocks:
with--image-mode segment - Embedded images:
with--image-mode embedded - Full page render:
with--image-mode page
- Segmented blocks:
Implementation Details
-
Text extraction and normalization
- Extracts selectable text from the PDF and collapses line breaks to form coherent paragraphs.
- Headings are inferred using font-size heuristics (larger font sizes are treated as heading markers).
-
Image extraction modes
segment(default): renders page segments/blocks to capture localized content (tables/figures) rather than entire pages.embedded: extracts images embedded in the PDF content stream.page: renders each full page as a single image (not recommended if you need cropped screenshots/blocks).
-
Filtering options
--filter-text on: runs OCR on extracted images and removes images whose OCR text density exceeds--text-threshold(e.g.,0.25).--filter-match on: removes images whose OCR text substantially matches the page's extracted text; controlled by--match-min-lenand language via--match-lang.--filter-pdf-text on: removes images that overlap PDF text blocks using PDF layout information (no OCR); controlled by--pdf-text-threshold(e.g.,0.1).
-
Output writing
- Writes a single Markdown file with
## Page XXsections and image links pointing to files saved under--image-dir.
- Writes a single Markdown file with
When Not to Use
- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
Required Inputs
- A clearly specified task goal aligned with the documented scope.
- All required files, identifiers, parameters, or environment variables before execution.
- Any domain constraints, formatting requirements, and expected output destination if applicable.
Recommended Workflow
- Validate the request against the skill boundary and confirm all required inputs are present.
- Select the documented execution path and prefer the simplest supported command or procedure.
- Produce the expected output using the documented file format, schema, or narrative structure.
- Run a final validation pass for completeness, consistency, and safety before returning the result.
Output Contract
- Return a structured deliverable that is directly usable without reformatting.
- If a file is produced, prefer a deterministic output name such as
pdf_extract_result.mdunless the skill documentation defines a better convention. - Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
Validation and Safety Rules
- Validate required inputs before execution and stop early when mandatory fields or files are missing.
- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
- Keep the output safe, reproducible, and within the documented scope at all times.
Failure Handling
- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
- If partial output is returned, label it clearly and identify which checks could not be completed.
Quick Validation
Run this minimal verification path before full execution when possible:
python scripts/extract_pdf.py --help
Expected output format:
Result file: pdf_extract_result.md
Validation summary: PASS/FAIL with brief notes
Assumptions: explicit list if any
Deterministic Output Rules
- Use the same section order for every supported request of this skill.
- Keep output field names stable and do not rename documented keys across examples.
- If a value is unavailable, emit an explicit placeholder instead of omitting the field.
Completion Checklist
- Confirm all required inputs were present and valid.
- Confirm the supported execution path completed without unresolved errors.
- Confirm the final deliverable matches the documented format exactly.
- Confirm assumptions, limitations, and warnings are surfaced explicitly.
Related Skills
Agent-Reach
86.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
84.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.5kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
