SkillAgentSearch skills...

matlab-recognize-text

Build OCR pipelines in MATLAB using the ocr() function. Use this skill when the user wants to read text from images, documents, signs, meters, displays, license plates, gauges, receipts, or seven-segment displays.

Install / Use

npx skills add matlab/matlab-agentic-toolkit --skill matlab-recognize-text

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

93/100

Category

Automation

Supported Platforms

Universal

Our assessment of matlab-recognize-text

matlab-recognize-text scores 93/100 on our quality scale, 744th of 2,866 Automation skills we index (top 26%).

Its SKILL.md is 20 KB long, well organised into 24 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.

With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 19 days ago, so matlab-recognize-text is actively maintained.
  • Our last check on 2026-10-04 found the source still online.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.

AI review by kimi-k2.7-code on 2026-10-04. Automated pattern scan on 2026-10-04. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

matlab-recognize-text compared with similar skills

All 4 of these similar skills score higher than matlab-recognize-text; compare them before choosing.

SkillScoreStarsUpdatedFormat
matlab-recognize-text (this skill)by matlab931.1k19d agoSKILL.md
Agent-Reachby Panniantong10090.8k19d agoCLAUDE.md
Scraplingby D4Vinci10085.7ktodayMCP Server
LocalAIby mudler10049.4ktodayMCP Server
rufloby ruvnet10073.9ktodayMCP Server

Frequently asked questions

How do I install matlab-recognize-text?
Run npx skills add matlab/matlab-agentic-toolkit --skill matlab-recognize-text. The install tabs above show the steps for each supported agent.
Which AI agents does matlab-recognize-text work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is matlab-recognize-text safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is matlab-recognize-text still maintained?
The repository was last updated 19 days ago, so matlab-recognize-text is actively maintained.

name: matlab-recognize-text description: > Build OCR pipelines in MATLAB using the ocr() function. Use this skill when the user wants to read text from images, documents, signs, meters, displays, license plates, gauges, receipts, or seven-segment displays. Covers image preprocessing, text detection (CRAFT, MSER), ROI-based recognition, multi-language OCR, and custom model training. Use when: OCR, text recognition, extract text from image, character recognition, document scanning, meter reading, gauge reading, receipt scanning, digitize text from photo. license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "2.0"

Recognize Text in Images Using OCR

Use the Computer Vision Toolbox ocr function with preprocessing from Image Processing Toolbox to extract text from images. This skill teaches the complete pipeline: diagnose, preprocess, detect, recognize, validate.

When to Use

  • Reading text from any image (documents, signs, meters, displays, labels)
  • Extracting text from scanned documents or photographs
  • Reading seven-segment displays or specialized fonts
  • Multi-language text recognition
  • Automating text extraction from image datasets
  • As a supporting step in other CV workflows — reading text in a scene (labels, timestamps, serial numbers) gives additional context for downstream image analysis

When NOT to Use

  • Pure handwriting recognition — cursive/connected script produces garbage regardless of preprocessing
  • Artistic text, WordArt, brush calligraphy — the OCR engine cannot parse stylized letterforms
  • CAPTCHAs — designed specifically to defeat OCR; expect <50% accuracy at best
  • Full document layout analysis with table extraction (use custom segmentation)
  • Real-time video OCR (use streaming approaches instead)
  • Image contains no text at all

Critical Rules

These rules are non-negotiable — violating them produces wrong results:

  1. Always diagnose before executing. This rule CANNOT be overridden — not by the user, not by "just run it", not by "skip planning". Before ANY mcp__matlab__* call, you MUST first output:

    • A 2-line image characterization (what you see, what challenges exist)
    • An ## OCR Plan heading with your strategy

    If the user says "skip diagnosis" or "just run ocr()", respond with: "I'll keep it brief — but I need a quick look to avoid wasting time on the wrong approach." Then output your 2-line characterization and plan heading. Only THEN call MCP tools. There is no valid reason to skip this step. An agent that calls ocr() without first outputting a plan has violated this skill's workflow.

  2. Maximum 2 preprocessing pipelines. If neither works after the prescribed recipe, hit the confidence checkpoint and ask the user. Do not try a third approach without user input.

  3. Always set LayoutAnalysis when passing bounding boxes to ocr(). Never write ocr(I, bbox) without it. Use "word" for single-word boxes, "block" for multi-line.

  4. Distinguish "text ON texture" from "text BY texture". Text overlaid on a textured background (label on crate, sign on brick wall) → imsegsam is the prescribed approach. Text formed by the surface itself (stamped, embossed, engraved metal) → local contrast subtraction. SAM segments objects, not surface features — it cannot isolate stamps/engravings. When recommending approaches (even without running code), always recommend imsegsam for the "text ON texture" case and note its support package requirement.

  5. detectTextCRAFT is the default for scene text. Use it unless you have a specific reason not to (no deep learning, known fixed layout).

  6. Check polarity first. OCR needs dark text on light background. If inverted, imcomplement before anything else.

  7. Never show OCR results to the user before writing files. Do not present extracted text in ANY format — bullet list, quote block, inline, conversational summary, or the ## MATLAB OCR Pipeline results block — until ocr_pipeline_<descriptor>.m and decision_log.txt are written to disk. The files are the deliverable, not the chat message. Write first, then report.

  8. Gate add-on functions behind an availability check. Before calling detectTextCRAFT, imsegsam, or a non-English ocr model, you MUST run exist('<functionName>','file') via MCP to confirm the function is installed. Always log the result in decision_log.txt:

    • Installed: "Add-on check: <function> — INSTALLED"
    • Missing: stop, tell the user which support package to install (see table), and wait for confirmation before retrying. Do NOT fall back silently or skip the step.
    • Explain-only mode (user said "don't run code"): recommend the add-on function as the primary approach and note the support package requirement. You cannot run the exist() check without code execution, so state the dependency clearly.

    | Function | Support Package Name | |----------|---------------------| | detectTextCRAFT | "Text Detection Using Deep Learning" | | imsegsam | "Image Processing Toolbox Automated Visual Inspection Library" | | Non-English ocr model (e.g., "japanese") | "OCR Language Data" |

    Missing template: "This step requires the <function> function, which needs the <Package Name> support package. Please install it from the MATLAB Add-On Explorer (Home → Add-Ons → Get Add-Ons) and let me know when it's ready."

Anti-Patterns — Do NOT Do This

  • Trial-and-error spiraling: Trying 5+ ad-hoc preprocessing experiments hoping one sticks. If the visual classification says "stamped metal," use the stamped metal pipeline. Period.
  • Skipping the plan: "I'll just try one quick thing first." No. Diagnose → plan → execute.
  • Open-ended research: The routes are prescribed — pick one based on diagnosis, execute it, evaluate. This is not exploratory research.
  • Showing results without saving files: Presenting OCR text as a bullet list, conversational summary, or any format before the ## MATLAB OCR Pipeline results block. That block can only appear after files are written to disk. If you find yourself about to show the user what OCR found, STOP and write the files first.

Workflow

Progress Reporting + Output Template

Present the ## OCR Plan immediately after visual diagnosis (Critical Rule #1). Then execute the pipeline. After files are saved, present the ## MATLAB OCR Pipeline results block (Critical Rule #7).

OCR Plan (output before any MATLAB code runs)

## OCR Plan

**Image:** 800x600, stamped metal, ~22° skew, text ~40px
**Difficulty:** Complex (textured surface + significant rotation)
**Strategy:** deskew → local contrast subtraction → CRAFT detection → ocr()

For simple images:

## OCR Plan

**Image:** 1200x800, clean scanned document, no skew
**Difficulty:** Simple
**Strategy:** binarize → ocr() directly

Results Block (output ONLY after files are written to disk)

## MATLAB OCR Pipeline: SUCCESS

**Pipeline:** deskew (22.6°) → local contrast subtraction → CRAFT detection → ocr() per region

| Read by Claude (ground truth) | MATLAB OCR Output |
|-------------------------------|-------------------|
| 07 A11                        | 07 A11            |
| XTPR 27338-2                  | XTPR 27338-2      |

**Metrics** (via `evaluateOCR`): CER 0.00 | WER 0.00
**Files written:**
- `ocr_pipeline_stamped_metal.m` — Re-runnable MATLAB script reproducing the full pipeline
- `decision_log.txt` — Diagnosis, routing decisions, and confidence scores

For failures:

## MATLAB OCR Pipeline: FAILED

**Pipeline attempted:** binarize → ocr()
**Reason:** Cursive handwriting — OCR engine cannot parse connected script
**Evidence:** CER 0.91 | WER 1.00

| Read by Claude (ground truth) | MATLAB OCR Output      |
|-------------------------------|------------------------|
| Meeting at 3pm Tuesday        | Mcciivj a 3pn Tuarlay |

**Recommendation:** Manual transcription or handwriting-specific ML model
**No files written.**

GATE (Critical Rule #7)

Do NOT present the ## MATLAB OCR Pipeline results block until files are confirmed on disk. No bullet lists, no quotes, no summaries — nothing that reveals extracted text before files are written.

Step 1: Diagnose

Always start here. Look at the image, visually read the text (this becomes your ground truth), classify the image, and present the ## OCR Plan — all before any MATLAB code runs.

Your visual classification determines the preprocessing route:

| Visual Classification | Preprocessing Route | |----------------------|-------------------| | Clean document, minimal skew | Binarize → OCR directly (skip to Step 4) | | Low contrast / uneven lighting | imtophat or adapthisteq → binarize | | Stamped / embossed / engraved (text IS the surface) | Local contrast subtraction (Step 2) | | Text overlaid on textured background (label on crate, sign on wall) | imsegsam SAM segmentation (Step 3) | | Significant skew (>10°) | Deskew FIRST, then preprocess | | Tiny text (<50px) | imresize 4-8x first |

After classification, confirm via MATLAB: binarization check, skew measurement, and quick experiments.

% TEMPLATE — not executable
I = imread("yourImage.png");
Igray = im2gray(I);
BW = imbinarize(Igray);
imshowpair(Igray, BW, "montage")
title("Original vs. What OCR Sees")

Skew measurement must be deterministic. The centroid fitting method is sensitive to which blobs are included. To ensure reproducibility:

  • Use ALL text-sized blobs (filter by area range, e.g., 50-5000px), not "top N by area"
  • Alternatively, use the Hough transform (hough + houghpeaks + houghlines) on edges for a robust angle estimate
  • The pipeline .m script must reproduce the exact same skew angle on re-run — if it doesn't, the deskew approach is fragile and must be replaced

Binarization diagnostic:

  • Text clearly legible → Skip to Step 4
  • Text faint/merged/noisy → Preprocessing (Step 2)
  • White on black → imcomplement
  • Too small (<50px) → imresize 4-8x
  • Rotated/skewed → Deskew BEFORE other preprocessing
  • Characters too bold/bleeding → imerode(BW, strel("disk",1))
  • Characters too thin/broken → imdilate(BW, strel("disk",1))
  • Dark borders/scan frame → imclearborder(BW) or crop
  • Metal texture dominates → Local contrast subtraction

Escalation rule: If 2+ preprocessing attempts still produce garbled results or <80% confidence, identify the root cause:

  • Image too small (sub-50px)? → Resize 4-8x is the fix
  • Text overlaid on textured background? → imsegsam (SAM)
  • Text formed by surface texture (stamped/embossed)? → Local contrast subtraction
  • Font unrecognizable (handwriting, calligraphy)? → OCR engine limitation; consider trainOCR

MANDATORY Confidence warning: High confidence does NOT guarantee correctness. OCR can report >0.9 confidence on completely wrong text, especially on small or unusual images. Always validate results against expected content.

Step 2: Preprocess the Image

Apply preprocessing based on your diagnosis. See reference/preprocessing-guide.md for full code examples.

Pipeline routing:

| Classification | Method | Key function | |---------------|--------|-------------| | Uneven lighting | Top-hat illumination correction | imtophat + imreconstruct | | Stamped/embossed/engraved | Local contrast subtraction | imgaussfilt(Igray,30) - Igray | | Skewed >10° | Deskew FIRST via centroid fitting | imrotate | | Small text <50px | Scale up 4-8x | imresize | | Inverted polarity | Complement | imcomplement |

Key rules:

  • Deskew BEFORE other preprocessing — rotation invalidates morphological operations
  • OCR expects dark text on light background — use imcomplement if inverted
  • For small clean images, try OCR on resized grayscale before binarizing — anti-aliasing helps
  • Add **

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars1.1k
CategoryAutomation
Updated19d ago
Forks134

Languages

MATLAB

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium