15 skills found
Huaizz-shawen / navigate-toOpenclaw + rosclaw + visual feedback
Huaizz-shawen / pick-objectOpenclaw + rosclaw + visual feedback
KuaaMU / mcp-vision-bridgeMCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.
sakamoto-family-smile / videodbSee, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips.
hassan-alnator / vibe-eyesAutomated visual testing framework for Claude and MCP-compatible clients. Generate unit tests, run E2E browser automation, perform visual regression testing, and extract text with OCR - all through natural language commands.
NVIDIA-AI-Blueprints / video-search-and-summarizationNVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting.
kitlau86 / agent-vision-mcpAn MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.
tengj / article-poster-generator将文章内容自动拆分为多张信息图海报(2400x3600竖版),支持多种风格模板和6种页面布局。 Use when: (1) 用户说"生成海报", (2) 用户发文章要求做图, (3) 用户说"做成卡片/信息图" (4) 用户发文章并希望转成可分享的图片, (5) 用户要求制作信息图/infographic
luanmorenommaciel / project-recapGenerate a visual HTML project recap — rebuild mental model of a project's current state, recent decisions, and cognitive debt hotspots
Huaizz-shawen / check-statusOpenclaw + rosclaw + visual feedback
oil-oil / grok-designerUse Grok 4.5 as the required external design advisor when a task needs UI critique, UX critique, component-choice review, interaction-flow review, design imagery markdown, art direction, visual hierarchy judgment, design-system fit, color/type/layout suggestions, HTML mockups, SVG icons, handwritten…
Shaohan-He / deepseek-eyes给 DeepSeek 装上眼睛 — MCP Server + 通义千问VL, 剪贴板图片→视觉模型→文字描述 / Give DeepSeek the ability to see images via clipboard + Qwen-VL
pierceboggan / nano-banana-mcpAn MCP server for generating images.
zombieyang / viteA Photoshop AI plugin
ProjectLiminality / DreamTalkA programmatic animation library extending the ancient modality of SandTalk into the digital domain