vision-skills
Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML file to image).
Install / Use
npx skills add Anionex/agent-vision-toolkit --skill vision-skillsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of vision-skills
vision-skills scores 93/100 on our quality scale, 776th of 4,626 Development & Engineering skills we index (top 17%).
Its SKILL.md is 15 KB long, well organised into 18 sections with 12 code examples: a thorough specification that gives an agent plenty to work with.
With 1,212 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 37 days ago, so vision-skills is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
vision-skills compared with similar skills
All 4 of these similar skills score higher than vision-skills; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| vision-skills (this skill)by Anionex | 93 | 1.2k | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 89.0k | 17d ago | CLAUDE.md |
| ai-job-searchby MadsLorentzen | 100 | 44.8k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 2d ago | CLAUDE.md |
| LocalAIby mudler | 100 | 49.4k | today | MCP Server |
Frequently asked questions
- How do I install vision-skills?
- Run
npx skills add Anionex/agent-vision-toolkit --skill vision-skills. The install tabs above show the steps for each supported agent. - Which AI agents does vision-skills work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is vision-skills safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is vision-skills still maintained?
- The repository was last updated 37 days ago, so vision-skills is actively maintained.
Skill content
View source on GitHubname: vision-skills description: >- Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operating a GUI from screenshots — and to re-check an image yourself when a description you were given lacks a detail.
vision-skills
Five local CLIs that give a text-only agent eyes. They read one shared
vision config (VISION_API_KEY / VISION_BASE_URL / VISION_MODEL /
LANG), plus the optional Python-client settings VISION_API_PROTOCOL,
VISION_REASONING_EFFORT, and VISION_USER_AGENT — no extra credentials.
Pick the tool by the question you are answering:
| Question | Tool |
|---|---|
| "What does this image show / say?" | glance |
| "Where is X?" — a thing you can name | ground |
| "Where are all the Xs?" — every instance of a kind | detect |
| "What is its exact shape, size, offset?" | trace |
| "Cut this box out as its own image file" | crop |
| "OCR this long screenshot / scrolling page / chat history" | scripts/long_screenshot_ocr.py |
| "Extract the icon/logo foreground as transparent PNG — manual region or auto (cropped+scaled screenshots)" | scripts/extract_fg.py |
| "Turn this HTML file into a viewport or full-page screenshot" | scripts/html_shot.py |
| "Which colours dominate a region, and which palette value fits it?" | scripts/dominant_colors.py |
| A relation none of them return — a gap, a distance between two located things | code over the pixels (Pillow) |
glance answers what something is; ground and detect answer where.
You give ground a description of a particular thing; you give detect a
kind and it enumerates the instances.
Both give real coordinates, but they are not pixel-exact: the box arrives
on a 0-1000 grid and is scaled to your image, so the last pixel or few are
not reliable. That is accurate enough to crop with, to click, to compare
positions against. When a number has to be exact, trace derives it from
the actual pixels — offsets, sizes, shapes.
Use the provided tools before hand-rolled pixels
Everything this toolkit ships a tool for, call the tool — do not rewrite it with Pillow in the middle of a task. The CLIs exist so the same pixel work is not hand-coded differently every time:
- cut a box out of an image →
crop, notImage.open(...).crop(...) - sample a region's palette →
scripts/dominant_colors.py - compare two images →
scripts/pixel_diff.py - vectorize to SVG →
trace - locate / inventory elements →
ground/detect - describe / OCR an image →
glance - safely split, OCR, and merge a long screenshot →
scripts/long_screenshot_ocr.py - HTML file to a viewport or full-page screenshot →
scripts/html_shot.py
Hand-written Pillow is only for what none of them return: a relation
between two things you already located (a gap, a distance), a resize or
overlay, drawing. If you catch yourself writing .crop(), .convert(),
or histogram code where one of the tools above fits, replace it with the
tool call — same coordinates, same box format, and the output feeds the
next tool directly.
glance — ask about an image
glance <image> # detailed description
glance <image> -q "<question>" # targeted question (qualitative only)
glance <image> --ocr # verbatim OCR
glance <image> --region X1,Y1,X2,Y2 -q "..." # zoom into a crop
glance <img1> <img2> -q "..." # compare in ONE call
When you do compare with glance, pass all paths to one call — separate
calls cannot see both images, so two descriptions compared afterwards are
two hallucination surfaces, not a comparison. --region uploads only the
crop, so small text and icons become readable.
But "what changed between these two?" is not a glance question. A one-word
badge or a small shift is a rounding error to a vision model and exact to
scripts/pixel_diff.py. Diff first to get the box, then glance --region
that box to read what the change actually is.
For a tall scrolling screenshot, do not send the whole image through one OCR
call and accept the model's downscaling loss. Run the long-screenshot workflow,
which finds low-content cut bands, invokes glance on each chunk, uses
structured extraction for chat histories, merges only duplicated overlap, and
writes a boundary audit:
python3 scripts/long_screenshot_ocr.py work/page.png -o work/page.ocr.md
python3 scripts/long_screenshot_ocr.py work/chat.png --mode chat --resume -o work/chat.ocr.md
Read references/long-screenshot-ocr.md before using it. It defines the
verification pass for unsafe cuts and chat-message boundaries.
ground — locate a named target
ground <image> "<target description>"
ground <image> "<target>" --region X1,Y1,X2,Y2
Output: x1: .., y1: .., x2: .., y2: .. in original-image pixels — with
--region too (crop hits are mapped back).
Provider-native 0-1000 boxes do not all use the same array order: Gemini uses
[y0, x0, y1, x1], while Qwen3-VL, Qwen3.5, and Qwen3.6 use
[x0, y0, x1, y1]. Grounding code must select the order by model family (or
an explicit override) before scaling to pixels; never parse every provider as
Gemini-style yxyx.
If several boxes come back numbered, your description matched more than one element rather than picking out a single thing. Narrow it with what distinguishes the one you mean — its text, its position, the block it sits in — and ask again.
The box is a handle, not just an answer — it feeds the next call:
$ ground screenshot.png "the send button"
x1: 1067, y1: 841, x2: 1108, y2: 881
$ glance screenshot.png --region 1067,841,1108,881 -q "is it enabled or greyed out?"
That two-step is how you inspect anything too small to survive a full-image pass.
detect — find every instance of a kind
detect <image> # every UI element
detect <image> "buttons" # one kind only
detect <image> --region X1,Y1,X2,Y2 # inside one box
You name a particular thing for ground; you name a kind for detect and
it enumerates the instances. Output is a numbered list with each item's
visible text and box. A full-screen
pass is a fast first draft — counts vary run to run on dense screens. For
completeness, detect the layout blocks first, then detect --region each
block.
trace — exact shape geometry (local, no vision API)
trace <image> # b/w spline SVG to stdout
trace <image> --polygon # boxy diagrams/wireframes
trace <image> --region X1,Y1,X2,Y2 -o out.svg # crop first
Coordinates come from the actual pixels, not a model's estimate. Flat,
high-contrast graphics only; text becomes curves (pair with --ocr when
the text matters). Small images are upscaled automatically before tracing,
so a 30px icon traces as readily as a screenshot — size is not a reason to
skip the tool. Before shipping or reusing a traced SVG, read
references/restore-graphic.md — it holds the reuse traps and the
ship-vs-hand-write call.
crop — cut a pixel box out of an image (local, no vision API)
crop <image> --region X1,Y1,X2,Y2 # writes <image-stem>.crop.png next to the input
crop <image> --region X1,Y1,X2,Y2 -o out.png
crop <image> --region X1,Y1,X2,Y2 --scale 4 # upscale the cut-out 4x (LANCZOS) first
The same X1,Y1,X2,Y2 pixel boxes ground/detect print, clamped to the
image bounds. Once a box is worth keeping — the same crop is about to feed
pixel_diff, dominant_colors, and trace in turn — cut it to a file
once and reuse it, instead of re-cropping in memory on every call.
--scale N upscales the cut-out before writing (default output name becomes
<image-stem>.crop@Nx.png): for icons too small for ground/trace to see
clearly, crop with --scale 4, then run ground/trace on the upscaled
file — coordinates it returns are in the upscaled grid, divide by N to map
back to the original image. Requires the optional pillow.
extract_fg — icon foreground as transparent PNG: manual region or auto (local, no vision API)
# manual: you know the region (and optionally the background colour)
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 -o icon.png
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --mode dark # grey/black line logos
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --exclude-color '#E6E6E6'
# auto: `crop --scale` cut-outs with the icon centred — no region needed
crop shot.png --region X1,Y1,X2,Y2 --scale 4 -o d/icon1.png
python3 scripts/extract_fg.py d/icon1.png d/icon2.png # writes <stem>.clean.png next to each input
python3 scripts/extract_fg.py d/icon1.png --disc-radius 60
python3 scripts/extract_fg.py d/icon1.png --boxes "101,84,184,171"
Manual mode keeps every sufficiently large connected component of the
region (separate logo sub-shapes stay together; specks drop out). Auto mode
takes a crop --scale cut-out with the icon centred (disc + glyph): the
disc centre is the image centre, the disc radius defaults to
min(w,h)/2 * 0.6, and the disc colour is sampled from a ring around the
centre; that colour is excluded and the glyph is picked as the most
saturated among the three largest coloured components (white rings,
ripples, and text fall away), output as a 1:1 transparent PNG. When auto
inference fails, override the radius with --disc-radius, or pass a
ground box (in the upscaled grid) as --boxes to recentre and re-filter
by overlap. Multiple images may be passed at once (auto mode).
Requires the optional pillow (and numpy for auto mode).
html_shot — render an HTML file to an image (local, needs a Chrome-family browser)
python3 scripts/html_shot.py page.html # writes page.png, 1280x800
python3 scripts/html_shot.py page.html --width 1440 --height 900 -o page.png
python3 scripts/html_shot.py page.html --scale 2 # 2x pixels: small text stays readable
python3 scripts/html_shot.py page.html --full-page # complete scroll height, same layout viewport
python3 scripts/html_shot.py page.html --full-page --max-pixels 40000000
The visual-alignment loop: write HTML, screenshot it at the reference
viewport, then compare it with the design. Use pixel_diff to locate
material differences, not to chase a zero-difference score. Rendering
happens in headless Chrome/Chromium/Edge — no Python dependencies. The default
captures only the viewport. Use --full-page for the complete document while
keeping --width and --height as the layout viewport, so vh/svh and
responsive breakpoints do not change. Add --max-pixels N when the page height
is untrusted. --wait-ms N pauses for fonts, images, or animation before
capturing. Paths are relative to this skill's own directory.
pixel_diff — where two images differ (local, no vision API)
python3 scripts/pixel_diff.py <a> <b> # path is relative to this skill dir
Prints an overall difference percentage plus the worst regions as x1: ..
boxes you can feed straight into glance --region. Exact where a vision
model rounds off.
dominant_colors — a region's palette, and the exact value among candidates (local, no vision API)
python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2 # top colour clusters + shares
python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2 \
--candidates '#F9FAFA,#F5F5F5,#F3F3F3,#EDEDED' # pick the best candidate
A vision model names a c
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
89.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ai-job-search
44.8kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
LocalAI
49.4kLocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
