vision-memory-mcp
Persistent visual cache for LLM-driven software development. Caches screenshots using perceptual hashing, vector search, and AX trees to prevent token overhead and visual hallucination loops.
Install / Use
claude mcp add putervision -- npx -y github:putervision/vision-memory-mcpIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of vision-memory-mcp
vision-memory-mcp scores 86/100 on our quality scale, 552nd of 956 AI & Machine Learning skills we index.
Its MCP Server is 7.1 KB long, well organised into 23 sections with 6 code examples: a thorough specification that gives an agent plenty to work with.
It has 49 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated yesterday, so vision-memory-mcp is actively maintained.
- Our last check on 2026-08-21 found the source still online.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 85/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
vision-memory-mcp compared with similar skills
All 4 of these similar skills score higher than vision-memory-mcp; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| vision-memory-mcp (this skill)by putervision | 86 | 49 | 1d ago | MCP Server |
| claude-memby thedotmack | 100 | 99.0k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 94.8k | 1d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.8k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.8k | today | CLAUDE.md |
Frequently asked questions
- How do I install vision-memory-mcp?
- Run
claude mcp add putervision -- npx -y github:putervision/vision-memory-mcp. The install tabs above show the steps for each supported agent. - Which AI agents does vision-memory-mcp work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is vision-memory-mcp safe to use?
- It declares no license and scores 85/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is vision-memory-mcp still maintained?
- The repository was last updated yesterday, so vision-memory-mcp is actively maintained.
Skill content
View source on GitHub@putervision/vision-memory-mcp
@putervision/vision-memory-mcp is a zero-infrastructure, local-first Model Context Protocol (MCP) server and CLI tool that provides AI coding assistants (such as Cursor, Claude Code, Gemini, or Copilot) with visual state caching using perceptual hashing, local CLIP embeddings, and transition graphs to eliminate repetitive vision LLM calls.
🌐 Official Documentation & Website: visionmemorymcp.com
⚡ Quick Start & Installation
Prerequisites: Node.js >= 18.17.0
1. Installation
# Global installation via npm
npm install -g @putervision/vision-memory-mcp
2. Workspace Initialization
Run init in your project root to scaffold database directories, .gitignore, .env, and IDE rules:
vision-memory-mcp init --yes
3. Basic MCP Client Setup
Add to your MCP client config (e.g. .cursor/mcp.json or .vscode/mcp.json):
{
"mcpServers": {
"vision-memory-mcp": {
"command": "vision-memory-mcp",
"args": ["run"]
}
}
}
Alternative Options & CLI Usage Examples
# Run stdio MCP server directly via binary (after global install)
vision-memory-mcp run
# Start server skipping heavy CLIP model downloads (air-gapped / offline mode)
vision-memory-mcp run --skip-model-load
# Re-initialize across all registered workspace projects
vision-memory-mcp init-global
# Health check dependencies, sharp bindings, and git safety
vision-memory-mcp doctor
# Run health diagnostics & aggregate metrics across all registered projects
vision-memory-mcp doctor-global
# Inspect stored visual states and metadata in terminal ASCII table
vision-memory-mcp inspect
# Register baseline design mockup contract (Visual SDD)
vision-memory-mcp spec set --name "Dashboard" --file ./dashboard-spec.png
# Save visual memory checkpoint snapshot
vision-memory-mcp snapshot save --name "v1.0-milestone"
# Ingest WebM / MP4 video recording into visual state memory timeline
vision-memory-mcp video ingest ./playwright-test.webm --category playwright_test
# Open interactive force-directed visual graph viewer in browser
vision-memory-mcp view
🌟 Key Highlights
- 👁️ Perceptual Visual Caching: Sub-5ms L1/L2 dHash zero-token fast-path layout recognition.
- 🎬 WebM & MP4 Video Ingestion: Digest E2E test recordings & screen captures into searchable keyframe visual states & state transition graphs.
- ⚡ 29 Core MCP Tools: Full visual state ingestion, video memory parsing, evidence packs (
create_evidence_pack), trajectory comparison, semantic vector retrieval, element grounding, visual SDD, and snapshot checkpoints. - 🔗 Dual-MCP Synergy & Immutable Evidence Packs: Deeply bridges
@putervision/state-memory-mcptask DAGs with visual state memory, generating cryptographically hashable evidence packs for compliance and audit trails. - 📉 Reduced Token Overhead: Caches UI states locally using dHash, local CLIP vector search, and accessibility trees so some savings on vision tokens can be expected.
- 🚀 Sub-5ms Fast-Path Latency: Eliminates repetitive vision LLM API calls and avoids visual hallucination loops.
- 🎯 Element Grounding & Action Target Prediction: Maps screen elements to CSS selectors and coordinates for deterministic UI interaction.
- 🎨 Visual Spec-Driven Development (Visual SDD): Register design mockups or screenshots as perceptual baseline contracts to verify visual regression.
- 🛡️ 100% Local-First Privacy: Local LanceDB vector store, local CLIP model, zero cloud telemetry, and PII redaction guarantees.
🚀 Architecture At a Glance
Incoming Screen
│
▼
┌──────────────────────────────┐
│ L1: In-Memory Cache Lookup │ ──(Hit)──▶ Return Cached Description & Grounded Elements
└──────────────┬───────────────┘
│ (Miss)
▼
┌──────────────────────────────┐
│ L2: Perceptual Hash Scan │ ──(Hit)──▶ Return Cached Description & Grounded Elements
└──────────────┬───────────────┘
│ (Miss)
▼
┌──────────────────────────────┐
│ L3: Local CLIP Vector Search │ ──(Hit)──▶ Return Semantically Close
└──────────────┬───────────────┘
│ (Miss)
▼
┌──────────────────────────────┐
│ L4: Vision LLM Fallback │ ──(Ingest)──▶ Save Redacted State to DB
└──────────────┬───────────────┘
📚 Documentation Directory
Explore dedicated guides and deep dives in the docs/ directory:
| Guide | Description |
| :--- | :--- |
| 🚀 Features & Architecture | Key features, 4-tier retrieval pipeline, element grounding, and Dual MCP Synergy. |
| 📘 Formal API Reference | Complete specifications, parameters, and schemas for all 23 MCP tools. |
| 🔌 Multi-IDE Integration Guide | Step-by-step configs for Cursor, Claude Desktop, Antigravity, Windsurf, Zed, Roo Code & Agent Rules. |
| 💻 CLI Commands Reference | Full guide for all 16 CLI management, visual spec, and snapshot commands. |
| ⚙️ Configuration Guide | Complete .env environment variables, thresholds, and L4 vision fallback setup. |
| 🔒 Storage Encryption & Security | Encryption details, local storage privacy, and PII masking guarantees. |
| 🤝 Contributing Guide | Development setup, codebase structure, and submission guidelines. |
| 🛡️ Security Policy | Security vulnerability reporting and privacy disclosures. |
| 📜 Changelog | Chronological record of release features, fixes, and patch updates. |
🧪 Testing
# Run full unit and integration test suite across all 37 test suites
npm run test
⚖️ License & Disclaimers
Developed and maintained by PuterVision LLC. Released under the MIT License.
- Local Storage Guarantee: Provided "as is" without warranty. Screenshots, perceptual hashes, vector embeddings, and transition graphs are stored locally unencrypted at the application level in
.vision-memory-mcp/. Zero telemetry or analytics data is ever transmitted. - Trademarks & Non-Affiliation: Product names (Cursor, Claude Code, Gemini, Windsurf, VS Code, Sharp, LanceDB, ONNX, HuggingFace) are property of their respective owners and used solely for compatibility identification.
Related Skills
claude-mem
99.0kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
94.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.8kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.8kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
