SkillAgentSearch skills...

yao-ocr

OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, bank cards, business licenses, tables, handwritten documents, or any visual text content.

Install / Use

npx skills add YaoApp/yao --skill yao-ocr

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

95/100

Supported Platforms

Universal

Our assessment of yao-ocr

yao-ocr scores 95/100 on our quality scale, 80th of 710 Content & Media skills we index (top 12%).

Its SKILL.md is 6.4 KB long, well organised into 14 sections with 10 code examples: a thorough specification that gives an agent plenty to work with.

With 8,024 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
29/30
Structure
20/20
Description
15/15
Adoption
17/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 4 days ago, so yao-ocr is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

yao-ocr compared with similar skills

All 4 of these similar skills score higher than yao-ocr; compare them before choosing.

SkillScoreStarsUpdatedFormat
yao-ocr (this skill)by YaoApp958.0k4d agoSKILL.md
Agent-Reachby Panniantong10085.8k12d agoCLAUDE.md
headroomby headroomlabs-ai10074.0k1d agoCLAUDE.md
crawl4aiby unclecode10084.4k3d agoMCP Server
Scraplingby D4Vinci10084.1ktodayMCP Server

Frequently asked questions

How do I install yao-ocr?
Run npx skills add YaoApp/yao --skill yao-ocr. The install tabs above show the steps for each supported agent.
Which AI agents does yao-ocr work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is yao-ocr safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is yao-ocr still maintained?
The repository was last updated 4 days ago, so yao-ocr is actively maintained.

name: yao-ocr description: OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, bank cards, business licenses, tables, handwritten documents, or any visual text content.

OCR Tools

Two tools for optical character recognition, supporting both VLM-OCR (vision language models) and traditional OCR APIs (Baidu, Google, Azure, PaddleOCR).

ocr_recognize

Extract text from images or PDF files using OCR.

Basic usage (plain text output):

tai tool ocr_recognize --source /path/to/image.png

With URL:

tai tool ocr_recognize --source https://example.com/document.jpg

Table extraction as Markdown:

tai tool ocr_recognize --source /path/to/table.png --type table --output_format markdown

Invoice structured extraction:

tai tool ocr_recognize --source /path/to/invoice.pdf --type invoice --output_format json

With specific provider:

tai tool ocr_recognize --source /path/to/doc.png --provider baidu

VLM-OCR with custom prompt:

tai tool ocr_recognize --source /path/to/doc.png --provider llm:qwen-ocr --prompt "只提取表格中的金额列"

PDF page range:

tai tool ocr_recognize --source /path/to/report.pdf --pages "1-5" --output_format markdown

| Parameter | Type | Required | Description | | ------------- | ------ | -------- | -------------------------------------------------------------------------------------------- | | source | string | yes | Image or PDF file path/URL to recognize | | provider | string | no | LLM connector ID (llm:xxx) or OCR settings key (baidu/paddleocr/google/azure). Auto-selects if omitted | | type | string | no | Recognition type (default: general). See type table below | | output_format | string | no | text (default), json (with coordinates/fields), or markdown (structured) | | mode | string | no | accurate (default, best quality) or standard (faster) | | language | string | no | Language hint (ISO 639-1, e.g. en, zh, ja). Auto-detected if omitted | | prompt | string | no | Custom instruction for VLM-OCR only, appended to system prompt. Ignored by traditional OCR | | pages | string | no | PDF page range, e.g. 1-5 or 1,3,7. All pages if omitted | | extra | JSON | no | Provider-specific parameters as a JSON object |

Recognition types

| Type | Description | Best output_format | | ---------------- | -------------------------- | ------------------ | | general | General text (default) | text | | table | Table extraction | markdown | | handwriting | Handwritten text | text | | document | Document layout parsing | markdown | | invoice | Invoice (VAT) | json | | receipt | Receipt / ticket | json | | id_card | ID card | json | | bank_card | Bank card | json | | license | Business license | json | | vehicle_license | Vehicle license | json | | passport | Passport | json | | license_plate | License plate | json |

If the chosen provider does not support the requested type, it automatically degrades to general and annotates the response metadata with degraded_from. VLM-OCR supports all types via prompt adaptation.

ocr_providers

List available OCR providers and their supported recognition types.

tai tool ocr_providers

Returns a list of providers including VLM-OCR models (from LLM connectors with ocr capability) and traditional API providers (from OCR settings). Each entry includes id, name, type (vlm or traditional), and supported_types.

PDF support

| Provider | PDF | Notes | | ---------- | --- | ---------------------------------------- | | Baidu | yes | pdf_file parameter | | Azure | yes | Document Intelligence native support | | PaddleOCR | yes | pdf + fileType=0 | | Google | no | Sync API does not support PDF | | VLM (llm:) | no | Vision models accept images only |

For providers that do not support PDF, use Baidu, Azure, or PaddleOCR instead.

Multi-page PDF response

Multi-page PDFs are automatically split page-by-page. Instead of printing all text, the tool returns a JSON summary with file paths for each page result:

{
  "source": "report.pdf",
  "total_pages": 10,
  "pages": 3,
  "results": [
    {"page": 1, "file": ".tool-tmp/ocr-a1b2c3d4/page-1.txt", "preview": "Invoice No: INV-001..."},
    {"page": 2, "file": ".tool-tmp/ocr-a1b2c3d4/page-2.txt", "preview": "Invoice No: INV-002..."},
    {"page": 3, "file": ".tool-tmp/ocr-a1b2c3d4/page-3.txt", "preview": "Invoice No: INV-003..."}
  ]
}

To read full content of a specific page, use cat:

cat .tool-tmp/ocr-a1b2c3d4/page-2.txt

Single-page PDFs and images return inline text as usual (no file indirection).

Guidelines

  • Use output_format=text (default) when you just need the text content — simplest for LLM processing
  • Use output_format=json for structured types (invoice, id_card, etc.) to get key-value fields
  • Use output_format=markdown for documents and tables to preserve layout
  • The prompt parameter only works with VLM-OCR providers; traditional OCR ignores it
  • For structured document types (invoice, receipt, id_card, etc.), prefer json output to get fields with key-value pairs
  • Use ocr_providers first to check which providers are available and what types they support
  • Multi-page PDFs return a JSON summary with temporary file paths; use cat <file> to read specific pages
  • Google Vision and VLM providers do not support PDF input directly; use Baidu, Azure, or PaddleOCR for PDF files

Related Skills

View on GitHub
GitHub Stars8.0k
CategoryContent
Updated4d ago
Forks716

Languages

Go

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium