SkillAgentSearch skills...

modlens

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

Install / Use

npx skills add liustack/modlens

Installs into whichever agent you are using.

About this skill
🤖

CLAUDE.md

Claude Code project instructions

Quality Score

95/100

Supported Platforms

Claude Code
OpenAI Codex

Our assessment of modlens

modlens scores 95/100 on our quality scale, 92nd of 963 AI & Machine Learning skills we index (top 10%).

Its CLAUDE.md is 18 KB long, well organised into 16 sections with 4 code examples: a thorough specification that gives an agent plenty to work with.

With 4,109 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so modlens is actively maintained.
  • Our last check on 2026-09-24 found the source still online.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands (1 minor note below). An AI review of the same text found nothing harmful.

  • noteInstalls by piping a downloaded script into a shellline 70
    curl -fsSL https://antigravity.google/cli/install.sh | bash

AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

modlens compared with similar skills

All 4 of these similar skills score higher than modlens; compare them before choosing.

SkillScoreStarsUpdatedFormat
modlens (this skill)by liustack954.1k5d agoCLAUDE.md
claude-memby thedotmack10095.3ktodayCLAUDE.md
Agent-Reachby Panniantong10089.5k18d agoCLAUDE.md
Understand-Anythingby Egonex-AI10085.2k1d agoCLAUDE.md
headroomby headroomlabs-ai10074.3ktodayCLAUDE.md

Frequently asked questions

How do I install modlens?
Run npx skills add liustack/modlens. The install tabs above show the steps for each supported agent.
Which AI agents does modlens work with?
It is written for Claude Code and OpenAI Codex, as a CLAUDE.md file. Other agents that read the same format can often use it too.
Is modlens safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands (1 minor note below). An AI review of the same text found nothing harmful. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is modlens still maintained?
The repository was last updated 5 days ago, so modlens is actively maintained.
<p align="center"> <img src="https://raw.githubusercontent.com/liustack/modlens/main/assets/banner.jpg" width="100%" alt="ModLens" /> </p> <h1 align="center">ModLens</h1> <p align="center"><b>Give a text-only model sight, and just paste the image.</b></p> <p align="center">🥇 <b>The most capable vision plugin for DeepSeek Harness (dsh)</b> 🥇</p> <p align="center"> <a href="./README.zh-CN.md">简体中文</a> · <a href="skills/modlens/references/configure.md">Configuration</a> · <a href="docs/troubleshooting.md">Troubleshooting</a> · <a href="docs/security.md">Security</a> · <a href="https://github.com/liustack/modsearch"><b>🔍 ModSearch (the best free web search plugin for DSH)</b></a> </p> <p align="center"> <a href="https://x.com/liustack"><img src="https://img.shields.io/badge/follow-%40liustack-black?style=flat-square&logo=x&logoColor=white" alt="Follow @liustack on X"></a> <a href="https://www.npmjs.com/package/@liustack/modlens"><img src="https://img.shields.io/npm/v/@liustack/modlens?style=flat-square&label=npm&color=cb3837" alt="npm"></a> <a href="https://nodejs.org"><img src="https://img.shields.io/node/v/@liustack/modlens?style=flat-square" alt="Node.js"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue?style=flat-square" alt="License"></a> <img src="https://img.shields.io/badge/Not%20backed%20by-Y%20Combinator-FF6600?style=flat-square&logo=ycombinator&logoColor=white" alt="Not backed by Y Combinator"> <img src="https://img.shields.io/badge/users-unknown-lightgrey?style=flat-square" alt="Users unknown"> </p>

DeepSeek's flagship chat models, and GLM-5.3 itself, are text-only and cannot read images. GLM-5.3-Flash is native multimodal. ModLens is a plug-in vision engine that gives a text-only model sight. ModLens reads images pasted straight into the chat, no saving to a file and passing a path first.

Talk to us

Issues are welcome any time: open one. Follow the liustack WeChat official account, and come find me on X: @liustack. What you built with it, which harness you are on, and what should come next are all shared on WeChat and X. A proper community space is on the way.

Highlights

🥇 The most capable vision plugin for DeepSeek Harness (dsh): install it instantly with one command: npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.26.5. See the setup guide for installation and update details. If the command line is not your thing but you still want to try DSH, check out <a href="https://github.com/liustack/aimanager"><b>AIManager</b></a>, the lightest desktop wrapper for DeepSeek Harness. It gets you started with zero code or configuration and installs every dependency for you with one click.

Pasting an image works two ways. ① Just paste. On a text-only model the pasted image lands as a private temp file and its path enters the composer (the same interaction OpenCode and Pi ship), then the modlens_read_image tool takes it from there. ② Pick a (modlens vision) entry in the model selector (it remembers your choice, so once is enough), then paste: the thumbnail stays visible in your message, closer to the Codex app feel, and the image is converted to structured evidence at request time, answered by the same underlying route. The plugin auto-discovers every provider route carrying eligible text-only DeepSeek, GLM, or MiMo Pro models and adds a wrapped entry per route. A stock install gets DeepSeek-V4-Flash (modlens vision) and DeepSeek-V4-Pro (modlens vision), while extra routes like opencode-go or zai get their own. Native vision models in those families, including GLM-5.3-Flash, are excluded automatically. Which paste route applies is the host's per-model call: only a model its metadata positively confirms text-only is taken over, anything unconfirmed is left alone, so vision models keep their native paste (details).

Paste images directly in every harness. No saving to a file and passing a path first.

A hotkey that captures the screen into DeepSeek Harness is a separate plugin: dsh-screenshot.

  • The lightest touch on the market. No hooks, no wrappers, no local proxy daemon, not a single line changed in any harness config: on the skill harnesses it is exactly one skill folder, on dsh exactly one plugin. Uninstalling is deleting a folder, and your agents are back to stock.
  • Zero-config start. Reuses existing setup in Claude Code, Codex, OpenCode, and Pi, plus other multimodal models already on your machine. Nothing installed locally? Antigravity CLI is a free no-key channel, and a free Gemini key brings a read down to 5-10 seconds. API keys from every major OpenAI-compatible provider work too.
  • Comma-separated keys rotate on auth, rate-limit, or quota failures. Other failures skip remaining keys and keep the existing provider failover.
  • Evidence, not imagination. Full transcription, reading-order layout regions, entity and relation lists. The model quotes specifics.
  • Install once, use everywhere. Verified on real machines in Claude Code, Codex, Pi, and OpenCode.

Install in other harnesses

Option 1, install with skills.sh:

npx -y skills add liustack/modlens --skill modlens --global

This installs the modlens skill at user level. Restart the harness, then ask your AI to configure modlens and run its health check.

Option 2, hand the install to your AI. Send it this line:

Install and configure the modlens skill following https://github.com/liustack/modlens/blob/main/INSTALL.md, then run the health check and tell me the result.

The install starts by checking what your machine already has. An existing login in Claude Code, Codex, OpenCode, or Pi can be enough: modlens asks before reusing any of them, and the health check tells you where things stand.

After either option, only if the health check comes back empty, set up a free engine. The recommended choice is a free Gemini API key (about three minutes at Google AI Studio, no credit card), which also makes every read 5-10 seconds. A free OpenAI-compatible key from another platform works too. To avoid any sign-up, install Antigravity CLI instead, then sign in:

curl -fsSL https://antigravity.google/cli/install.sh | bash
agy                                                           # sign in, then exit

The install also inventories vision reachable through your other local harness CLIs (Codex, OpenCode, Pi) and asks, per harness, whether modlens may reuse it. Granted logins join the engine pool as equals, and every reused read is labeled with whose quota it spent.

On DeepSeek Harness the command line is not the only way in. Settings → Plugins → Plugin config carries a ModLens card (from dsh 0.1.7 on, open the sidebar's Plugins page and pick @liustack/modlens under Installed): switch the engine, tick which local CLIs auto mode may reuse, hit save and it takes effect.

The ModLens vision-engine card in the dsh settings page, shown in Chinese: switch the engine, tick which local CLIs auto mode reuses

Usage

Once installed, just chat. Paste an image or drop a path, ask anything, and the skill triggers on its own: the image goes to a vision engine and the answer comes back grounded in what it read. Paste once, and later questions about the same image do not need another paste.

Vision engines: six built-in providers, four reusable CLIs, one failover chain

ModLens does not depend on any single vision service. Ten sources of vision in total: six built-in providers, any one of which is enough, plus four local agent CLIs whose logins can be reused. The built-ins:

| Provider | What it needs | Speed per read | Good for | | :-- | :-- | :-- | :-- | | gemini-api | a free Gemini API key (3 minutes, no card) | 5-10s | the recommended default | | openai | any OpenAI-compatible endpoint (key + baseUrl + model) | 5-10s | qwen-vl, GLM, self-hosted gateways | | anthropic | an Anthropic API key | 5-10s | machines already holding one | | antigravity-cli | the free agy CLI, one browser sign-in, no key | 15-45s | zero-signup starts | | claude-cli | a signed-in Claude Code | 20-45s | riding your existing Claude subscription | | kimi-cli | a signed-in Kimi Code | 20-45s | riding your existing Kimi subscription, named explicitly |

Without a pinned provider, every configured engine forms one failover chain: the fast API providers try first, the agent CLIs back them up, the first good result wins, and meta.attempts records every attempt so a fallback is never silent.

openai is a universal socket, not just OpenAI

Any endpoint speaking the OpenAI chat-completions protocol with image input plugs straight in — that covers most of the vision-model world:

modlens config set openai.baseUrl https://dashscope.aliyuncs.com/compatible-mode/v1   # qwen-vl
modlens config set openai.apiKey  <key>
modlens config set openai.model   qwen3-vl-plus

apiKey (and the matching env var) also accepts a comma-separated list. ModLens rotates to the next key after authentication, rate-limit, or quota failures. Network, 5xx, and parse failures skip remaining keys and keep provider failover.

The same three keys work for GLM's open platform, SiliconFlow, OpenRouter, a self-hosted vLLM/Ollama, or any gateway of your own. If your favorite vision model has an OpenAI-compatible API, ModLens can drive it.

Reusing what your machine already has

Two more sources of vision need zero new keys, each behind one explicit consent recorded in config:

  • The harness you are talking in right now. Running inside Claude Code with a subscription signed in? claude-cli reads images through it out of the box. The install flow asks the same question for whichever harness you install into.
  • Every other agent CLI on the machine. modlens doctor discovers them, you grant per harness, and they join the same failover chain with no priority over your own keys. Every reused read is labeled in meta.warnings with whose quota it spent, so nothing is ever silently billed:

| Reused CLI | What it needs | Grant with | Rides as | | :-- | :-- | :-- | :-- | | Codex | a signed-in Codex CLI with a vision model | config set reuse.codex true | agent lane, 15-45s | | OpenCode | a vision model configured in OpenCode | config set reuse.opencode true | agent lane, 15-45s | | Pi | model credentials held by Pi | config set reuse.pi true | an API key upgrades to the 5-10s inline lane, OAuth drives Pi itself | | Grok | a signed-in Grok CLI (SuperGrok) | config set reuse.grok true | agent lane, 15-45s |

Picking and routing

Two knobs: modlens config set provider <name> states a preference (the chain still backs it up), -p <name> pins exactly one with no fallback. Machines behind a proxy set HTTPS_PROXY or modlens config set proxy <url> and the API providers route through it. An internal endpoint can opt out with modlens config set openai.proxy "". Details: the CLI manual for defaults and flags, Configuration for every key, and Security for who fetches what on remote URLs.

See it work

Unedited runs, all driving a text-only DeepSeek-V4-Flash.

The newest one first: pasting a screenshot straight into DeepSeek Harness on the DeepSeek-V4-Flash (modlens vision) variant. The paste keeps its native thumbnail, the trajectory shows the image arriving "already transcribed by the modlens vision bridge", and the answer walks the UI element by element.

![Pasting an image straight into DeepSeek Harness, read through the mod

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars4.1k
CategoryAI
Updated5d ago
Forks128

Languages

TypeScript

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions