genome
Auditable memory layer for AI agents: zero-LLM-call local ingest (~10ms/msg, air-gapped), matches Mem0 on accuracy at ~1000x lower ingest cost, bi-temporal belief-state, MCP server. Honest LoCoMo/LongMemEval benchmarks. Open source (Apache-2.0).
Install / Use
claude mcp add NORTHTEKDevs -- npx -y github:NORTHTEKDevs/genomeIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of genome
genome scores 77/100 on our quality scale, 817th of 955 AI & Machine Learning skills we index.
Its MCP Server is 22 KB long, well organised into 27 sections with 22 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 6 days ago, so genome is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
genome compared with similar skills
All 4 of these similar skills score higher than genome; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| genome (this skill)by NORTHTEKDevs | 77 | 10 | 6d ago | MCP Server |
| claude-memby thedotmack | 100 | 98.1k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 93.9k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.6k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.7k | today | CLAUDE.md |
Frequently asked questions
- How do I install genome?
- Run
claude mcp add NORTHTEKDevs -- npx -y github:NORTHTEKDevs/genome. The install tabs above show the steps for each supported agent. - Which AI agents does genome work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is genome safe to use?
- It is Apache-2.0-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is genome still maintained?
- The repository was last updated 6 days ago, so genome is actively maintained.
Skill content
View source on GitHubGENOME
Open memory for AI agents. Same answer accuracy as Mem0 - but ~1,000× cheaper to store, runs fully offline, and keeps an auditable record.
Papers: Do Agents Need an LLM to Remember? (the core evaluation, 2026) and What Does Each Memory Feature Buy? (a measured audit of all five optional features, wins and failures alike, 2026). PDFs in papers/; result tables in benchmarks/AUDIT-RESULTS.md.
Most agent-memory tools (like Mem0) call an LLM on every message to decide what to remember. That's the slow, expensive part - and GENOME's bet is that you don't need it. GENOME just embeds each message locally: no LLM, no API, no network in the write path.
Benchmarked honestly on public datasets (LoCoMo, LongMemEval), GENOME answers just as accurately as Mem0 - while storing memories for a tiny fraction of the cost and running completely offline.
Honest up front: on answer accuracy, GENOME ties Mem0 - we do not claim to beat it there (six independent benchmark configurations confirm parity, none significant in either direction). The advantage is cost, speed, offline operation, and a temporal/auditable record Mem0 can't produce.
See it work

Every frame is real output from examples/demo_timeline.py,
captured by tools/render_demo_gif.py. Run it yourself,
no API key required:
python examples/demo_timeline.py
The interesting part is step 3. The same question gets three different correct answers depending on when you ask about, because the store keeps when each fact became true rather than overwriting it:
| Question | Answer | |---|---| | What was Priya's city in May 2023? | Boston [Mar 2023 - Jan 2024] | | What was Priya's city in March 2024? | Seattle [Jan 2024 - Feb 2025] | | What is Priya's city now? | Austin [Feb 2025 - present] |
The "thinking about maybe moving to Denver, nothing decided" turn is stored but never becomes an answer: it is a plan, not a durable fact.
How it works
The write path is deliberately dumb and cheap. All the intelligence happens at read time, when there is a query to focus it.
flowchart LR
M["incoming message"] --> E["local embedder<br/>all-MiniLM-L6-v2"]
E --> S[("local store<br/>SQLite or Postgres")]
M -. "optional, opt-in" .-> B["belief extraction<br/>(the only LLM call)"]
B --> K[("bi-temporal<br/>fact log")]
Q["query"] --> R["exact cosine search<br/>over this tenant's rows"]
S --> R
R --> RR["optional cross-encoder<br/>rerank"]
RR --> A["context for the agent"]
Q --> PIT["as-of resolution<br/>facts_valid_at(entity, T)"]
K --> PIT
PIT --> A
style E fill:#0A84FF,color:#fff
style S fill:#1c2530,color:#fff
style K fill:#1c2530,color:#fff
style B fill:#3a3a3a,color:#fff
Write: embed locally, store. About 10 ms, zero LLM calls, zero network calls. The embedding is deterministic -- the same text always yields the same vector, with no sampled extraction step deciding what matters -- so what gets stored is a function of the input, and replaying a journal reproduces that store exactly. (Ids and timestamps are stamped per write, so two independent ingests of the same conversation agree on content and vectors, not on record ids.)
Read: exact cosine search within the tenant's scope (no ANN index to build or update), with an optional local cross-encoder reranker.
Bi-temporal layer (opt-in): records each fact at its domain time, the moment it became true in the world, not the moment it was ingested. That is what makes point-in-time questions answerable even when facts arrive out of order.
Why the record can be re-derived
flowchart TB
subgraph LLM["LLM-extraction memory"]
A1["message"] --> A2["LLM decides what matters<br/>(sampled, non-deterministic)"]
A2 --> A3[("store")]
A3 --> A4["replaying the same input<br/>can produce a different store"]
end
subgraph GEN["GENOME"]
B1["message"] --> B2["local embedding<br/>(deterministic)"]
B2 --> B3[("store")]
B3 --> B4["replaying the same input<br/>reproduces the same store"]
end
style A4 fill:#5c1f1f,color:#fff
style B4 fill:#1f4d33,color:#fff
A record that cannot be re-derived is difficult to audit. That property, not accuracy, is the actual argument for this design.
Don't believe it? Prove it yourself
The cost, speed, and offline claims need no API key - measure them on your machine in 60 seconds:
git clone https://github.com/NORTHTEKDevs/genome && cd genome
pip install -e . && python -m genome.verify
The first run downloads the local embedding model (~90 MB, one time) before printing anything, so expect 30-120 seconds of apparent silence on a cold machine. Every run after that is instant.
It writes memories with your outbound network physically blocked and prints a live pass/fail receipt - 0 network calls, 0 LLM calls, single-digit-ms writes, retrieval that works:
[PASS] Air-gapped write path: wrote 200 memories with every outbound socket blocked -> 0 network attempts, 0 LLM calls
[PASS] Write latency: 7.1 ms/message (Mem0's measured write path: ~2,055 ms + 1 LLM call/message)
[PASS] Retrieval works: top hit score 0.598
That receipt covers the cost/speed/offline story only. The accuracy-parity with Mem0 claim
is a separate, larger check that needs an LLM key - reproduce it head-to-head on the same
questions with your own key via python benchmarks/head_to_head.py (one OpenRouter key works;
see benchmarks/RESULTS.md for the n=90 / n=205 runs, the paired
significance tests, and the published nulls). The full test suite runs in public CI (badge
above). The pitch isn't "trust me" - it's "run it."
Add persistent memory to your agent in one line (MCP)
GENOME ships a fully-local MCP server - cross-session memory for Claude Desktop, Claude Code, or Cursor with no API key and no data leaving your machine:
pip install "genome-memory[mcp]"
{ "mcpServers": { "genome": { "command": "genome-mcp" } } }
Or zero-install via uv: { "command": "uvx", "args": ["--from", "genome-memory[mcp]", "genome-mcp"] }
Tools the agent gets: remember, recall, forget, reset_memories.
Memories persist locally in ~/.genome/memories.db. Full MCP details ↓
GENOME vs Mem0 at a glance
| | GENOME | Mem0 | |---|---|---| | Answer accuracy (LoCoMo, LongMemEval) | tied | tied | | LLM calls to store one message | 0 | 1+ | | Write speed | ~10 ms | ~2,000 ms | | Runs offline / air-gapped | yes | no (needs an LLM API) | | Ingest cost (10k-user deployment) | ~$190 / yr | $159k-$1.6M / yr | | "What was true in March?" (point-in-time) | yes | no | | Deterministic, auditable memory | yes | no |
Every number is measured within one harness - same responder, judge, embedder, and top-k;
only the memory layer changes - with paired significance tests. Full detail and per-number
provenance: benchmarks/RESULTS.md. Formatted report:
benchmarks/GENOME-LoCoMo-Report.pdf.
Why it's ~1,000× cheaper: it never calls an LLM to remember
Storing one message costs one LLM call in Mem0, zero in GENOME (just a local embedding). That's not a benchmark you can argue with - it's arithmetic, and it holds no matter which LLM you price it against. At 10,000 users × 50 messages/day (15M messages/month):
| Model Mem0 uses to extract | Mem0's yearly ingest bill | GENOME | |---|---|---| | Claude Haiku | $1,601,757 | $190 | | gpt-4o-mini | $238,596 | $190 | | cheapest hosted model | $159,064 | $190 |
The gap survives the cheapest model and grows in production (Mem0 re-sends stored memories
to the LLM as the store fills). Reproduce: python benchmarks/tco_project.py (no API key).
It runs air-gapped
GENOME's default embedder is local. We proved the write path is genuinely offline by blocking all network during writes - they still succeed:
- ~10 ms/message, 0 network calls, 0 LLM calls (
python benchmarks/local_writepath.py) - Mem0 can't do this - it needs an LLM API call to ingest.
That makes GENOME usable on-prem, in regulated environments, or fully offline. It's a yes/no capability, not a price point.
How it works
- Write: embed the message locally and store it. No LLM, no network. (~10 ms)
- Read: vector search over your memories, with an optional local cross-encoder reranker for harder queries.
- Optional bi-temporal layer: track how facts change over time and answer "what was true at time T" - see below.
What determinism buys you
Because nothing on the write path interprets your content, GENOME can do things an LLM-ingest memory system cannot do in principle:
-
Memory firewall (
genome.firewall): tag every write with where it came from (user,agent,tool,web), quarantine low-trust origins from recall, and enforce origin-bound authority - web content can never UPDATE or DELETE what your user said, even when a prompt-injected conflict resolver asks for it. There is also no extraction step for injected content to attack: the write path has no LLM.from genome import Memory from genome.firewall import TrustPolicy m = Memory(trust_policy=TrustPolicy(recall_min_trust=1)) m.add("I live in Anchorage", user_id="u1", provenance="user") m.add(scraped_page_text, user_id="u1", provenance="web") # quarantined -
Explainable recall (
genome.explain):explain_search()reports every candidate's dense score, BM25 rank, fused score, and - when it was not returned - the exact reason (parent-filtered, quarantined, beyond the limit). Two runs agree, so a recall bug can be committed as a regression test instead of a shrug. -
Journal + replay (
genome.journal): record every mutation and provably reproduce the store -verify_journal()replays the history and compares canonical hashes. Replay a prefix to roll back; replay into different storage to branch a memory for a what-if run. The journal sits after extraction, so replay is deterministic even if you configured an LLM extractor. Each line chains to its predecessor, so a removed or edited line is detected even when the change cancels out in the final state.# Tamper-EVIDENT by default. Pass a key (kept outside the journal's directory) # to make it tamper-PROOF: an unkeyed chain can be recomputed by anyone with # write access, an HMAC chain cannot. m = Memory(journal="mem.journal", journal_key=os.environb[b"GENOME_JOURNAL_KEY"]) -
Multi-agent belief attribution (
record_fact(..., believed_by="agent-a")): agents sharing a store keep their own belief timelines - agent B disagreeing does not clobber agent A's fact - and
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
98.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
93.9kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.6kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.7kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
