mem0-server-mcp
π§ Production-ready MCP server providing intelligent memory for Claude Code with async architecture, Neo4j knowledge graphs, smart chunking & enterprise security. One-command Docker deployment.
Install / Use
claude mcp add subhashdasyam -- npx -y github:subhashdasyam/mem0-server-mcpIf the server publishes to npm under a different name, use that package instead β check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
SecuritySupported Platforms
Skill content
View source on GitHubπ§ Mem0 MCP Server - Self-Hosted Memory for AI
A production-ready, self-hosted Model Context Protocol (MCP) server that provides persistent, intelligent memory for Claude Code and other AI assistants. Features async/await architecture, knowledge graph intelligence, smart text chunking, and enterprise-grade security. Built with Docker Compose for one-command deployment.
β¨ Features
Core Features
- π One-Command Deployment - Start the entire stack with a single script
- π 100% Self-Hosted - No external API dependencies (when using Ollama)
- π Token-Based Authentication - Secure multi-user access with PostgreSQL-backed token management
- π Multi-LLM Support - Works with Ollama, OpenAI, or Anthropic
- π― Project Isolation - Automatic memory isolation per project directory
- π Semantic Search - Vector-based search with pgvector
- β‘ 13 MCP Tools - Complete memory management + intelligence analysis
- π Dual Transport Support - Modern HTTP Stream (recommended) + legacy SSE transport
- π³ Docker Compose - Easy orchestration of all services
- π§ͺ Comprehensive Tests - Automated test suite included
- π Audit Logging - Track all authentication attempts and token usage
π§ Memory Intelligence System
- π Knowledge Graphs - Link memories with typed relationships (RELATES_TO, DEPENDS_ON, SUPERSEDES, etc.)
- π Temporal Tracking - Track how knowledge evolves over time
- ποΈ Architecture Mapping - Map system components and dependencies
- π Impact Analysis - Understand cascading effects of changes
- π Decision Tracking - Record technical decisions with pros/cons/alternatives
- π― Topic Clustering - Automatically detect knowledge groups
- β Quality Scoring - Trust scores based on validations and citations
- π Intelligence Analysis - Comprehensive health reports with actionable recommendations
π¦ Smart Text Chunking System
- βοΈ Semantic Chunking - Automatically splits large text at paragraph/sentence boundaries
- π Context Preservation - 150-character overlap between chunks maintains context continuity
- β‘ Performance Optimization - Prevents timeouts on large text inputs with 8B+ embedding models
- π·οΈ Chunk Metadata - Full tracking with chunk index, total chunks, size, and overlap indicators
- π Session Continuity - All chunks share the same
run_idfor related memory grouping - π― Transparent Operation - Small texts (<1000 chars) bypass chunking for optimal performance
ποΈ Architecture
ββββββββββββββββ
β Claude Code β (Your IDE with MCP client)
ββββββββ¬ββββββββ
β HTTP Stream (recommended): http://localhost:8080/mcp
β SSE (legacy): http://localhost:8080/sse
β + Token Authentication Headers
β
ββββββββββββββββ
β MCP Server β Port 8080 (FastMCP)
β (Python) β β’ 13 MCP Tools (5 core + 8 intelligence)
β β β’ Token Validation
β β β’ Dual Transport Support
ββββββββ¬ββββββββ
β HTTP REST API
β
ββββββββββββββββ
β Mem0 Server β Port 8000 (FastAPI)
β (FastAPI) β β’ 28 REST Endpoints (13 core + 15 intelligence)
β β β’ Multi-LLM Support
β β β’ Vector + Graph Storage
β β β’ Memory Intelligence System
β β β’ Async/Await Architecture with Background Tasks
ββββββββ¬ββββββββ
β
βββββ΄βββββ¬βββββββββββ¬βββββββ
β β β β
ββββββββββ ββββββ βββββββ ββββββββ
βPostgresβ βNeo4jβ βAuth β βOllamaβ
βpgvectorβ βGraphβ βTokenβ β LLM β
βVector β βIntel-β βStoreβ β β
βSearch β βligenceβ β β β β
ββββββββββ ββββββ βββββββ ββββββββ
β‘ Async Architecture
The Mem0 server uses FastAPI's async/await architecture for optimal performance:
- Non-blocking I/O: Handles multiple requests concurrently without blocking
- Background Neo4j Sync: Memories stored immediately in PostgreSQL, then synced to Neo4j asynchronously
- Retry Logic: Automatic retry with exponential backoff (7 attempts: 1s, 2s, 4s, 8s, 16s, 32s)
- Immediate Response: API returns instantly without waiting for graph sync
- Fault Tolerance: If Neo4j sync fails, memory still accessible via PostgreSQL vector search
This architecture ensures fast response times even when processing complex graph operations.
π Quick Start (5 Minutes)
Prerequisites
-
Docker & Docker Compose installed
docker --version docker compose version -
Ollama Server with models (or OpenAI/Anthropic API key)
# On your Ollama server: ollama pull qwen3:8b ollama pull qwen3-embedding:8b
Installation
# 1. Clone or copy this directory
cd /path/to/mem0-mcp
# 2. Create configuration file
cp .env.example .env
# 3. Edit .env with your Ollama server address (if needed)
nano .env # Update OLLAMA_BASE_URL
# 4. Start everything!
./scripts/start.sh
That's it! The script will:
- β Start PostgreSQL with pgvector
- β Start Neo4j graph database
- β Start Mem0 REST API server
- β Start MCP server for Claude Code
- β Wait for all services to be healthy
Setup Authentication
Step 1: Run Database Migrations
./scripts/migrate-auth.sh
Step 2: Create Your Authentication Token
python3 scripts/mcp-token.py create \
--user-id your.email@company.com \
--name "Your Name" \
--email your.email@company.com
This will output your token and setup instructions. Copy the MEM0_TOKEN value.
Step 3: Add to Your Shell Profile
Add these lines to ~/.zshrc or ~/.bashrc:
export MEM0_TOKEN='mcp_abc123...' # Your token from step 2
export MEM0_USER_ID='your.email@company.com'
Then reload:
source ~/.zshrc # or ~/.bashrc
Connect to Claude Code
Recommended: Using Claude CLI (Easiest)
# Add mem0 server with HTTP Stream transport (recommended)
claude mcp add mem0 http://localhost:8080/mcp/ -t http \
-H "X-MCP-Token: ${MEM0_TOKEN}" \
-H "X-MCP-UserID: ${MEM0_USER_ID}"
# Verify it's configured
claude mcp list
Alternative: Manual Configuration
Add this to your Claude Code MCP configuration file:
File: ~/.config/claude-code/config.json
{
"mcpServers": {
"mem0": {
"url": "http://localhost:8080/mcp/",
"transport": "http",
"headers": {
"X-MCP-Token": "your-token-here",
"X-MCP-UserID": "your.email@company.com"
}
}
}
}
Legacy SSE Transport (Backward Compatibility)
# Using CLI
claude mcp add mem0 http://localhost:8080/sse/ -t http \
-H "X-MCP-Token: ${MEM0_TOKEN}" \
-H "X-MCP-UserID: ${MEM0_USER_ID}"
Important:
- Always include the trailing slash in URLs:
/mcp/or/sse/(not/mcpor/sse) - HTTP Stream transport (
/mcp/) is recommended as it's the modern MCP protocol - SSE (
/sse/) is maintained for backward compatibility
Restart Claude Code and you're ready to go!
π Usage
Basic Commands
# Start the stack
./scripts/start.sh
# View logs
./scripts/logs.sh # All services
./scripts/logs.sh mem0 # Specific service
# Check health
./scripts/health.sh
# Run tests
./scripts/test.sh
# Stop the stack
./scripts/stop.sh
# Restart
./scripts/restart.sh
# Clean all data (β οΈ destructive)
./scripts/clean.sh
Using with Claude Code
Once connected, you can use these commands in Claude Code:
"Store this code in memory: [your code snippet]"
"Search my memories for Python functions"
"Show all my stored coding preferences"
"Delete memory with ID [id]"
"Show history of memory [id]"
Available MCP Tools
Core Memory Tools (5)
- add_coding_preference - Store code snippets and implementation details
- search_coding_preferences - Semantic search through memories
- get_all_coding_preferences - Retrieve all stored memories
- delete_memory - Delete specific memory by ID
- get_memory_history - View change history
Memory Intelligence Tools (8)
- link_memories - Create typed relationships between memories (build knowledge graphs)
- get_related_memories - Graph traversal to discover connected context
- analyze_memory_intelligence π - GAME-CHANGER: Comprehensive intelligence report with health scores, clusters, and recommendations
- create_component - Map system architecture with component nodes
- link_component_dependency - Define dependencies between components
- analyze_component_impact - Analyze cascading effects of changes
- create_decision - Track technical decisions with pros/cons/alternatives
- get_decision_rationale - Retrieve decision context and reasoning
All tools automatically use authentication credentials from your MCP configuration headers.
Using Memory Intelligence
"Link these two memories as related"
"Show me all memories related to authentication"
"Analyze my project's knowledge graph health"
"Create a component called Database with type Infrastructure"
"What would be impacted if I change the Authentication component?"
"Record this decision: Use PostgreSQL, pros: ACID compliance, cons: complexity"
π Authentication Management
Token Management
# List all tokens
python3 scripts/mcp-token.py list
# List tokens for specific user
python3 scripts/mcp-token.py list --user-id john.doe@company.com
# Create a new token
python3 scripts/mcp-token.py create \
--user-id john.doe@company.com \
--name "John Doe" \
--email john.doe@company.com
# Revoke (disable) a token
python3 scripts/mcp-token.py revoke mcp_abc123
# Re-enable a token
python3 scripts/mcp-token.py enable mcp_abc123
# Delete a token permanently
python3 scripts/mcp-token.py delete mcp_abc123
# View audit log
python3 scripts/mcp-token.py audit --days 30
# View user statistics
python3 scripts/mcp-token.py stats john.doe@company.com
Testing Authentication
./tests/test_auth.sh
This tests: missing headers, invalid tokens, user ID mismatches, valid authentication, token revocation, and re-enabling.
π§ Configuration
LLM Providers
Ollama (Default - Free, Self-Hosted)
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://192.168.1.2:11434
OLLAMA_LLM_MODEL=qwen3:8b
OLLAMA_EMBEDDING_MODEL=qwen3-embedding:8b
OLLAMA_EMBEDDING_DIMS=4096
Supported Models:
- LLM: llama3, qwen3, mistral, phi3, etc.
- Embeddings: qwen3-embedding (4096d), nomic-embed-text (768d), all-minilm (384d)
OpenAI (Cloud - Paid)
LLM_PROVIDER=openai
OPENAI_API_KEY=sk-proj-...
OPENAI_LLM_MODEL=gpt-4o
OPENAI_EMBEDDING_MODEL=text-embedding-3-small
OPENAI_EMBEDDING_DIMS=1536
Anthropic (Cloud - Paid, LLM only)
LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
ANTHROPIC_MODEL=claude-3-5-sonnet-20241022
# Still need embeddings from Ollama or OpenAI:
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_EMBEDDING_DIMS=768
Project Isolation
Control how memories are isolated per project:
# Auto mode (recommended) - Auto-detect project from directory
PROJECT_ID_MODE=auto
# Manual mode - Set explicitly per project
PROJECT_ID_MODE=manual
DEFAULT_USER_ID=my_project_name
# Global mode - Share all memories
PROJECT_ID_MODE=global
DEFAULT_USER_ID=shared_memory
Performance Tuning
For faster performance, use smaller embedding models:
# Fast: 768 dimensions (enables HNSW indexing)
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_EMBEDDING_DIMS=768
# Slower but more accurate: 4096 dimensions (HNSW disabled)
OLLAMA_EMBEDDING_MODEL=qwen3-embedding:8b
OLLAMA_EMBEDDING_DIMS=4096
Note: pgvector's HNSW index is limited to
Truncated for display β read the full file on GitHub.
Related Skills
momen-cursurrules-prompt-file
40.6kCursor rules for building custom frontends with Momen.app as headless BaaS with GraphQL API, actionflows, AI agents, and Stripe integration.
semiotic-react-dataviz-cursorrules-prompt-file
40.6kCursor rules for Semiotic data visualization library with 30+ chart types, MCP server, and AI-assisted chart generation.
Agent-Reach
71.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu β one CLI, zero API fees.
ruflo
67.7kπ The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
