open-dream-rsi
Recursive Self-Improvement through Evolving Worlds (Dream-RSI) for LLM agents
Install / Use
claude mcp add patrykorwat -- npx -y github:patrykorwat/open-dream-rsiIf the server publishes to npm under a different name, use that package instead โ check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of open-dream-rsi
open-dream-rsi scores 72/100 on our quality scale, 883rd of 966 AI & Machine Learning skills we index.
Its MCP Server is 39 KB long, well organised into 45 sections with 28 code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated today, so open-dream-rsi is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit โ read the skill file before letting an agent act on it.
open-dream-rsi compared with similar skills
All 4 of these similar skills score higher than open-dream-rsi; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| open-dream-rsi (this skill)by patrykorwat | 72 | 10 | today | MCP Server |
| claude-memby thedotmack | 100 | 99.7k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 95.9k | 3d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.9k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 75.1k | today | CLAUDE.md |
Frequently asked questions
- How do I install open-dream-rsi?
- Run
claude mcp add patrykorwat -- npx -y github:patrykorwat/open-dream-rsi. The install tabs above show the steps for each supported agent. - Which AI agents does open-dream-rsi work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is open-dream-rsi safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is open-dream-rsi still maintained?
- The repository was last updated today, so open-dream-rsi is actively maintained.
Skill content
View source on GitHubOpen Dream-RSI
๐ Preprint: Open Dream-RSI: An Open-Source Library for Recursive Self-Improvement Around a Frozen LLM, with Replay-Gated Learned Policies and a Curated Knowledge Base โ P. Orwat, 2026. PDF ยท LaTeX source
What is this?
Every AI agent pays the same tax over and over: it fails a task, works out what went wrong, and then throws that understanding away when the session ends. Open Dream-RSI keeps what was learned โ without touching the model.
The LLM stays frozen. What improves is everything around it:
- the exploration policy (even re-written as Python code by the LLM itself, gated by replay before it is trusted),
- a library of verified solutions (recipes) reused as warm starts,
- curated lessons distilled from failures โ every one of them promoted only after it measurably helps on replay,
- and the world model the policy is scored against: a Discovery Tree of every attempt, enriched with replay-safe error-class facts by the Sentinel engine.
Between runs the system "dreams": it replays thousands of strategies over the recorded history in-process, at zero API cost, and keeps only what beats the incumbent. An open implementation of Recursive Self-Improvement through Evolving Worlds (Dream-RSI, Zheng et al., arXiv:2609.14858).
Zero runtime dependencies โ Python 3.10+ stdlib only, ~10.2k lines of library code, ~5.1k lines of tests, no build step.
Not another memory plugin. Experience-memory systems record what the agent learned and reuse it as-is. ODR is an evolving worlds loop: every attempt is recorded into a Discovery Tree, and between runs the system dreams โ replaying thousands of strategy variants against that recorded world. Nothing is promoted on the model's word alone: a candidate policy, recipe or lesson replaces the incumbent only after surviving counterfactual replay.
Evaluation policy: this project makes performance claims on exactly
one benchmark โ goose ร TravelPlanner, public and scored by the
benchmark's own evaluators (below).
The scripted suites (bench, bench-policy) are deterministic
self-checks โ behavioural regression contracts for the machinery โ and
carry no model-quality claim.
How it works
โโโโโโโโโโโโโโโโโ
โ Frozen LLM โ
โโโโโโโโฌโโโโโโโโโ
โ proposals
online loop โผ offline dream
โโโโโโโโโโโโโ DiscoveryTree โโโโโโโโโโโโ ReplaySimulator โโโโโโโโโโโ
โ โ โ โ
โ node.sentinel counterfactual โ
โ (Sentinel facts, rollout scoring โ
โ replay-safe) โ โ
โ โ โผ โ
โ โโโโโโ Policy improvement โโโโโ next cycle โ
โ โ
SENTINEL MODE (tool-execution layer) โ
โ โ โ โ
error-class nudge permute-gate โโโ block tool calls โ
ledger (once/session) past 80% of budget โโโโโโโโโโโโ
- Online execution โ the agent attempts each task with the real LLM and a sandboxed verifier, appending every attempt to a per-task Discovery Tree (code, test feedback, score, one-line plan, and โ when a Sentinel engine is wired โ structured error-class facts and an explicit termination reason).
- Offline dreaming โ instead of paying for real-world rollouts, the agent "dreams" over the recorded history: thousands of strategy variants are scored in an in-process replay simulator at zero external-call cost. The best parameters, programs and recipes are persisted and steer the next cycle.
Beyond parameter-level dreaming, the loop closes section 3 of the paper
("dreaming with code"): each cycle the LLM may rewrite the exploration
policy itself as a small Python program, choose_action(frontier, step).
Candidates are statically validated (AST gate), executed only in a hardened
subprocess sandbox, and scored by counterfactual replay rollout on the
recorded tree. A candidate replaces the incumbent only on evidence โ and
the incumbent is re-scored on the current tree, not trusted at its stored
score. A crashing or cheating policy can never break the loop: expansion
falls back to the greedy baseline.
Quick start
git clone https://github.com/patrykorwat/open-dream-rsi.git
cd open-dream-rsi
pip install -e .
Point it at a model โ or don't. Any OpenAI-compatible endpoint works;
keys are read only from the environment. The model name is
auto-detected from the endpoint (GET /v1/models), so pointing at a local
vLLM/Ollama needs only the base URL:
export OPENAI_BASE_URL=http://127.0.0.1:8000/v1 # vLLM / Ollama / any gateway
export OPENAI_API_KEY=*** # any token for local servers
# no ODR_LLM_MODEL needed: the first model the endpoint serves is used.
# Pin ODR_LLM_MODEL only to select among several served models.
Queue a task in tasks.json:
[{
"task_id": "add1",
"category": "math",
"prompt": "Implement add(a, b) returning the sum.",
"tests": [{"call": "add(2, 3)", "expected": 5}],
"max_attempts": 3
}]
Run the supervisor โ once, as a daemon, or as the recommended automated
dream cycle (one command that imports session evidence, runs the promotion
gate when the evidence justifies it, and publishes earned lessons as skills):
python -m open_dream_rsi --memory .dream_rsi dream \
--sessions ~/.hermes/state.db --skills-out ~/.hermes/skills/odr-curated
python -m open_dream_rsi loop --tasks tasks.json --interval 300 # daemon mode
python -m open_dream_rsi loop --tasks tasks.json --once # single cycle
python -m open_dream_rsi status # what it learned
No key? Everything deterministic still runs: bench, bench-policy,
gate-replay and the dashboard all use a scripted mock client. A full
in-process walkthrough with a mock OpenAI server: examples/live_loop_demo.py.
Don't want the CLI at all? The same loop runs as an MCP server inside your existing coding agent โ that is the fastest way in:
python3 -m open_dream_rsi mcp --tasks ./tasks.json --memory ./.dream_rsi
โฆsee the next section for the exact block for your agent.
Plug in your favourite agent
The loop speaks MCP (Model Context Protocol). One config block and your everyday agent can queue tasks for the dreamer, pull back verified recipes and consult curated lessons โ zero extra LLM setup: the dreamer resolves its own brain (env โ local goose config โ localhost vLLM), and asks the endpoint itself which model to use.
| Tool | What it does |
|---|---|
| odr_status | what the loop has learned (policies, recipes, events) |
| odr_recipes | best verified solution for a category โ use as warm start |
| odr_lessons | curated failure lessons for a category / search query |
| odr_add_task | queue a task (prompt + tests, or prompt + criteria) โ test-less tasks are completion-judged |
| odr_run_once | run one improvement cycle now (bounded API budget) |
| odr_dream | the automated cycle: import host session evidence, distill curated lessons, run the promotion gate, publish skills |
Tools are registered prefixed per client (e.g. Hermes
mcp_open_dream_rsi_odr_run_once). No model calls tools on its own
initiative โ trigger them via a recipe's instructions:, a hook, or an
explicit ask ("check odr_recipes before you try again"). A tip that makes
the loop proactive: add to your project AGENTS.md:
## Self-improvement loop (MCP: open-dream-rsi)
- Before implementing a self-contained Python utility, call `odr_recipes`
and `odr_lessons` for its category; reuse a verified recipe verbatim.
- When you finish a hard, testable function, queue it with `odr_add_task`
(task_id = function name, tests = your test cases) so the loop can
dream over it offline.
OpenCode
Create (or merge into) opencode.json in the project root where you
work:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"open-dream-rsi": {
"type": "local",
"command": ["python3", "-m", "open_dream_rsi", "mcp",
"--tasks", "./tasks.json", "--memory", "./.dream_rsi"],
"enabled": true
}
}
}
Start OpenCode and type /mcps โ you should see open-dream-rsi: connected.
Goose
Recommended โ the script wires everything (diagnoses the install, resolves
your goose provider, optionally starts the model-borrowing proxy, and
writes the extension block into ~/.config/goose/config.yaml itself):
./scripts/odr_goose_setup.sh # diagnose + configure goose + start proxy
./scripts/odr_goose_setup.sh --check # diagnose only, change nothing
./scripts/odr_goose_setup.sh --no-proxy # skip the proxy (use provider='goose')
Or add the extension manually to ~/.config/goose/config.yaml:
extensions:
open-dream-rsi:
enabled: true
type: stdio
name: open-dream-rsi
description: "Dream-RSI self-improvement loop. Use odr_add_task to queue a
Python task with tests, odr_run_once to run an improvement cycle,
odr_recipes/odr_lessons to reuse verified solutions and failure lessons,
odr_status to inspect what the loop has learned."
cmd: python3
args: ["-m", "open_dream_rsi", "mcp",
"--tasks", "/ABSOLUTE/PATH/projects/myproject/tasks.json",
"--memory", "/ABSOLUTE/PATH/projects/myproject/.dream_rsi"]
timeout: 300
Goose spawns extensions without a shell and with a scrubbed environment:
use absolute paths (no ~, no relative ./) and pass any env the server
needs via envs: โ it will not inherit your export'ed OPENAI_*. After
any config edit: fully quit goose (the desktop app caches config at
startup), relaunch, and activate open-dream-rsi in the session's
extensions picker. One-shot alternative (no config edit):
goose session --with-extension "python3 -m open_dream_rsi mcp --tasks ./tasks.json --memory ./.dream_rsi".
Two first-class ways to run the dreamer on exactly the model goose uses
(no second API key): provider: "goose" reads ~/.config/goose/config.yaml
itself (CLI and desktop dialects, custom_providers/*.json, secrets.yaml,
macOS keychain), and python3 -m open_dream_rsi proxy --port 8799 exposes
goose's brain as a plain OpenAI-compatible endpoint โ details in
docs/integrations.md.
Hermes Agent
Add under mcp_servers in ~/.hermes/config.yaml (or via the dashboard's
MCP catalog):
mcp_servers:
open-dream-rsi:
command: "python3"
args: ["-m", "open_dream_rsi", "mcp",
"--tasks", "/ABS/PATH/open-dream-rsi/tasks.json",
"--memory", "/ABS/PATH/open-dream-rsi/.dream_rsi"]
timeout: 300
env:
OPENAI_BASE_URL: "http://YOUR-VLLM-HOST:8000/v1"
OPENAI_API_KEY: "***" # any token for local servers
Restart Hermes; tools appear as mcp_open_dream_rsi_* in every platform
toolset. Hermes is also the one host with MCP sampling support, if you
ever want the dreamer to borrow
Truncated for display โ read the full file on GitHub.
Related Skills
claude-mem
99.7kPersistent Context Across Sessions for Every Agent โ Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
95.9kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu โ one CLI, zero API fees.
Understand-Anything
85.9kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
75.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit โ see the Safety scan above for what the skill file itself contains.
