18 skills found
kitlau86 / agent-vision-mcpAn MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.
KuaaMU / mcp-vision-bridgeMCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.
symgraph / BinAssistMCPBinary Ninja plugin to provide MCP functionality.
NVIDIA-AI-Blueprints / video-search-and-summarizationNVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting.
Shaohan-He / deepseek-eyes给 DeepSeek 装上眼睛 — MCP Server + 通义千问VL, 剪贴板图片→视觉模型→文字描述 / Give DeepSeek the ability to see images via clipboard + Qwen-VL
pierceboggan / nano-banana-mcpAn MCP server for generating images.
Dataojitori / nocturne_memoryA lightweight, rollbackable, and visual Long-Term Memory Server for MCP Agents. Say goodbye to Vector RAG and amnesia. Empower your AI with persistent, graph-like structured memory across any model, session, or tool. Drop-in replacement for OpenClaw.
jordanrendric / claude-video-visionGive Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis
merterbak / Grok-MCPMCP server for xAI’s Grok API with Web/X search, vision, image/video generation and file support
MCPBlender / blender-mcp🎨 Control Blender 3D with Claude AI — prompt-driven 3D modeling, materials & scene generation via MCP
manascb1344 / together-mcp-serverMCP server enabling high-quality image generation via Together AI's Flux.1 Schnell model.
Huaizz-shawen / navigate-toOpenclaw + rosclaw + visual feedback
cedricvidal / memorious-mcpSemantic Memory for MCP. 100% Local & Private. Store, recall, and forget with vector search via ChromaDB
wuyoscar / GPT-Image2-SkillGPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing
huang-sh / pytorch-fsdpAn open-source, model-agnostic AI workbench for scientific discovery.
hassan-alnator / vibe-eyesAutomated visual testing framework for Claude and MCP-compatible clients. Generate unit tests, run E2E browser automation, perform visual regression testing, and extract text with OCR - all through natural language commands.
Huaizz-shawen / check-statusOpenclaw + rosclaw + visual feedback
sakamoto-family-smile / videodbSee, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips.