SkillAgentSearch skills...

mem0-server-mcp

🧠 Production-ready MCP server providing intelligent memory for Claude Code with async architecture, Neo4j knowledge graphs, smart chunking & enterprise security. One-command Docker deployment.

Install / Use

claude mcp add subhashdasyam -- npx -y github:subhashdasyam/mem0-server-mcp

If the server publishes to npm under a different name, use that package instead β€” check the repo README.

About this skill
πŸ”Œ

MCP Server

Model Context Protocol server

Quality Score

70/100

Category

Security

Supported Platforms

Claude Code
Claude Desktop

🧠 Mem0 MCP Server - Self-Hosted Memory for AI

A production-ready, self-hosted Model Context Protocol (MCP) server that provides persistent, intelligent memory for Claude Code and other AI assistants. Features async/await architecture, knowledge graph intelligence, smart text chunking, and enterprise-grade security. Built with Docker Compose for one-command deployment.

License: MIT Docker Python

✨ Features

Core Features

  • πŸš€ One-Command Deployment - Start the entire stack with a single script
  • πŸ”’ 100% Self-Hosted - No external API dependencies (when using Ollama)
  • πŸ” Token-Based Authentication - Secure multi-user access with PostgreSQL-backed token management
  • 🌐 Multi-LLM Support - Works with Ollama, OpenAI, or Anthropic
  • 🎯 Project Isolation - Automatic memory isolation per project directory
  • πŸ“Š Semantic Search - Vector-based search with pgvector
  • ⚑ 13 MCP Tools - Complete memory management + intelligence analysis
  • πŸ”Œ Dual Transport Support - Modern HTTP Stream (recommended) + legacy SSE transport
  • 🐳 Docker Compose - Easy orchestration of all services
  • πŸ§ͺ Comprehensive Tests - Automated test suite included
  • πŸ“ Audit Logging - Track all authentication attempts and token usage

🧠 Memory Intelligence System

  • πŸ”— Knowledge Graphs - Link memories with typed relationships (RELATES_TO, DEPENDS_ON, SUPERSEDES, etc.)
  • πŸ•’ Temporal Tracking - Track how knowledge evolves over time
  • πŸ—οΈ Architecture Mapping - Map system components and dependencies
  • πŸ“Š Impact Analysis - Understand cascading effects of changes
  • πŸ“ Decision Tracking - Record technical decisions with pros/cons/alternatives
  • 🎯 Topic Clustering - Automatically detect knowledge groups
  • ⭐ Quality Scoring - Trust scores based on validations and citations
  • πŸš€ Intelligence Analysis - Comprehensive health reports with actionable recommendations

πŸ“¦ Smart Text Chunking System

  • βœ‚οΈ Semantic Chunking - Automatically splits large text at paragraph/sentence boundaries
  • πŸ”„ Context Preservation - 150-character overlap between chunks maintains context continuity
  • ⚑ Performance Optimization - Prevents timeouts on large text inputs with 8B+ embedding models
  • 🏷️ Chunk Metadata - Full tracking with chunk index, total chunks, size, and overlap indicators
  • πŸ”— Session Continuity - All chunks share the same run_id for related memory grouping
  • 🎯 Transparent Operation - Small texts (<1000 chars) bypass chunking for optimal performance

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Claude Code β”‚  (Your IDE with MCP client)
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚ HTTP Stream (recommended): http://localhost:8080/mcp
       β”‚ SSE (legacy): http://localhost:8080/sse
       β”‚ + Token Authentication Headers
       ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  MCP Server  β”‚  Port 8080 (FastMCP)
β”‚  (Python)    β”‚  β€’ 13 MCP Tools (5 core + 8 intelligence)
β”‚              β”‚  β€’ Token Validation
β”‚              β”‚  β€’ Dual Transport Support
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚ HTTP REST API
       ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Mem0 Server  β”‚  Port 8000 (FastAPI)
β”‚  (FastAPI)   β”‚  β€’ 28 REST Endpoints (13 core + 15 intelligence)
β”‚              β”‚  β€’ Multi-LLM Support
β”‚              β”‚  β€’ Vector + Graph Storage
β”‚              β”‚  β€’ Memory Intelligence System
β”‚              β”‚  β€’ Async/Await Architecture with Background Tasks
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
   β”Œβ”€β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”
   ↓        ↓          ↓      ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”
β”‚Postgresβ”‚ β”‚Neo4jβ”‚  β”‚Auth β”‚ β”‚Ollamaβ”‚
β”‚pgvectorβ”‚ β”‚Graphβ”‚  β”‚Tokenβ”‚ β”‚ LLM  β”‚
β”‚Vector  β”‚ β”‚Intel-β”‚  β”‚Storeβ”‚ β”‚      β”‚
β”‚Search  β”‚ β”‚ligenceβ”‚ β”‚     β”‚ β”‚      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜

⚑ Async Architecture

The Mem0 server uses FastAPI's async/await architecture for optimal performance:

  • Non-blocking I/O: Handles multiple requests concurrently without blocking
  • Background Neo4j Sync: Memories stored immediately in PostgreSQL, then synced to Neo4j asynchronously
  • Retry Logic: Automatic retry with exponential backoff (7 attempts: 1s, 2s, 4s, 8s, 16s, 32s)
  • Immediate Response: API returns instantly without waiting for graph sync
  • Fault Tolerance: If Neo4j sync fails, memory still accessible via PostgreSQL vector search

This architecture ensures fast response times even when processing complex graph operations.

πŸš€ Quick Start (5 Minutes)

Prerequisites

  1. Docker & Docker Compose installed

    docker --version
    docker compose version
    
  2. Ollama Server with models (or OpenAI/Anthropic API key)

    # On your Ollama server:
    ollama pull qwen3:8b
    ollama pull qwen3-embedding:8b
    

Installation

# 1. Clone or copy this directory
cd /path/to/mem0-mcp

# 2. Create configuration file
cp .env.example .env

# 3. Edit .env with your Ollama server address (if needed)
nano .env  # Update OLLAMA_BASE_URL

# 4. Start everything!
./scripts/start.sh

That's it! The script will:

  • βœ… Start PostgreSQL with pgvector
  • βœ… Start Neo4j graph database
  • βœ… Start Mem0 REST API server
  • βœ… Start MCP server for Claude Code
  • βœ… Wait for all services to be healthy

Setup Authentication

Step 1: Run Database Migrations

./scripts/migrate-auth.sh

Step 2: Create Your Authentication Token

python3 scripts/mcp-token.py create \
  --user-id your.email@company.com \
  --name "Your Name" \
  --email your.email@company.com

This will output your token and setup instructions. Copy the MEM0_TOKEN value.

Step 3: Add to Your Shell Profile

Add these lines to ~/.zshrc or ~/.bashrc:

export MEM0_TOKEN='mcp_abc123...'  # Your token from step 2
export MEM0_USER_ID='your.email@company.com'

Then reload:

source ~/.zshrc  # or ~/.bashrc

Connect to Claude Code

Recommended: Using Claude CLI (Easiest)

# Add mem0 server with HTTP Stream transport (recommended)
claude mcp add mem0 http://localhost:8080/mcp/ -t http \
  -H "X-MCP-Token: ${MEM0_TOKEN}" \
  -H "X-MCP-UserID: ${MEM0_USER_ID}"

# Verify it's configured
claude mcp list

Alternative: Manual Configuration

Add this to your Claude Code MCP configuration file:

File: ~/.config/claude-code/config.json

{
  "mcpServers": {
    "mem0": {
      "url": "http://localhost:8080/mcp/",
      "transport": "http",
      "headers": {
        "X-MCP-Token": "your-token-here",
        "X-MCP-UserID": "your.email@company.com"
      }
    }
  }
}

Legacy SSE Transport (Backward Compatibility)

# Using CLI
claude mcp add mem0 http://localhost:8080/sse/ -t http \
  -H "X-MCP-Token: ${MEM0_TOKEN}" \
  -H "X-MCP-UserID: ${MEM0_USER_ID}"

Important:

  • Always include the trailing slash in URLs: /mcp/ or /sse/ (not /mcp or /sse)
  • HTTP Stream transport (/mcp/) is recommended as it's the modern MCP protocol
  • SSE (/sse/) is maintained for backward compatibility

Restart Claude Code and you're ready to go!

πŸ“– Usage

Basic Commands

# Start the stack
./scripts/start.sh

# View logs
./scripts/logs.sh          # All services
./scripts/logs.sh mem0     # Specific service

# Check health
./scripts/health.sh

# Run tests
./scripts/test.sh

# Stop the stack
./scripts/stop.sh

# Restart
./scripts/restart.sh

# Clean all data (⚠️  destructive)
./scripts/clean.sh

Using with Claude Code

Once connected, you can use these commands in Claude Code:

"Store this code in memory: [your code snippet]"
"Search my memories for Python functions"
"Show all my stored coding preferences"
"Delete memory with ID [id]"
"Show history of memory [id]"

Available MCP Tools

Core Memory Tools (5)

  1. add_coding_preference - Store code snippets and implementation details
  2. search_coding_preferences - Semantic search through memories
  3. get_all_coding_preferences - Retrieve all stored memories
  4. delete_memory - Delete specific memory by ID
  5. get_memory_history - View change history

Memory Intelligence Tools (8)

  1. link_memories - Create typed relationships between memories (build knowledge graphs)
  2. get_related_memories - Graph traversal to discover connected context
  3. analyze_memory_intelligence πŸš€ - GAME-CHANGER: Comprehensive intelligence report with health scores, clusters, and recommendations
  4. create_component - Map system architecture with component nodes
  5. link_component_dependency - Define dependencies between components
  6. analyze_component_impact - Analyze cascading effects of changes
  7. create_decision - Track technical decisions with pros/cons/alternatives
  8. get_decision_rationale - Retrieve decision context and reasoning

All tools automatically use authentication credentials from your MCP configuration headers.

Using Memory Intelligence

"Link these two memories as related"
"Show me all memories related to authentication"
"Analyze my project's knowledge graph health"
"Create a component called Database with type Infrastructure"
"What would be impacted if I change the Authentication component?"
"Record this decision: Use PostgreSQL, pros: ACID compliance, cons: complexity"

πŸ” Authentication Management

Token Management

# List all tokens
python3 scripts/mcp-token.py list

# List tokens for specific user
python3 scripts/mcp-token.py list --user-id john.doe@company.com

# Create a new token
python3 scripts/mcp-token.py create \
  --user-id john.doe@company.com \
  --name "John Doe" \
  --email john.doe@company.com

# Revoke (disable) a token
python3 scripts/mcp-token.py revoke mcp_abc123

# Re-enable a token
python3 scripts/mcp-token.py enable mcp_abc123

# Delete a token permanently
python3 scripts/mcp-token.py delete mcp_abc123

# View audit log
python3 scripts/mcp-token.py audit --days 30

# View user statistics
python3 scripts/mcp-token.py stats john.doe@company.com

Testing Authentication

./tests/test_auth.sh

This tests: missing headers, invalid tokens, user ID mismatches, valid authentication, token revocation, and re-enabling.

πŸ”§ Configuration

LLM Providers

Ollama (Default - Free, Self-Hosted)

LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://192.168.1.2:11434
OLLAMA_LLM_MODEL=qwen3:8b
OLLAMA_EMBEDDING_MODEL=qwen3-embedding:8b
OLLAMA_EMBEDDING_DIMS=4096

Supported Models:

  • LLM: llama3, qwen3, mistral, phi3, etc.
  • Embeddings: qwen3-embedding (4096d), nomic-embed-text (768d), all-minilm (384d)

OpenAI (Cloud - Paid)

LLM_PROVIDER=openai
OPENAI_API_KEY=sk-proj-...
OPENAI_LLM_MODEL=gpt-4o
OPENAI_EMBEDDING_MODEL=text-embedding-3-small
OPENAI_EMBEDDING_DIMS=1536

Anthropic (Cloud - Paid, LLM only)

LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
ANTHROPIC_MODEL=claude-3-5-sonnet-20241022

# Still need embeddings from Ollama or OpenAI:
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_EMBEDDING_DIMS=768

Project Isolation

Control how memories are isolated per project:

# Auto mode (recommended) - Auto-detect project from directory
PROJECT_ID_MODE=auto

# Manual mode - Set explicitly per project
PROJECT_ID_MODE=manual
DEFAULT_USER_ID=my_project_name

# Global mode - Share all memories
PROJECT_ID_MODE=global
DEFAULT_USER_ID=shared_memory

Performance Tuning

For faster performance, use smaller embedding models:

# Fast: 768 dimensions (enables HNSW indexing)
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_EMBEDDING_DIMS=768

# Slower but more accurate: 4096 dimensions (HNSW disabled)
OLLAMA_EMBEDDING_MODEL=qwen3-embedding:8b
OLLAMA_EMBEDDING_DIMS=4096

Note: pgvector's HNSW index is limited to

Truncated for display β€” read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars9
CategorySecurity
Updated10mo ago
Forks1

Languages

Python

Security Score

86/100

Audited on Oct 13, 2025

2 low