SkillAgentSearch skills...

agent-vision-mcp

An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.

Install / Use

claude mcp add kitlau86 -- npx -y github:kitlau86/agent-vision-mcp

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

33/100

Supported Platforms

Claude Code
Claude Desktop
Gemini CLI

Related Skills

View on GitHub
GitHub Stars10
CategoryAI
Updated11h ago
Forks2

Languages

TypeScript

Security Score

97/100

Audited on Aug 8, 2026

1 info