aizynthfinder-retrosynthesis
AiZynthFinder retrosynthetic route planning (CASP) from AstraZeneca Molecular AI. Monte Carlo tree search guided by a template-based neural expansion policy recursively disconnects a target SMILES until precursors are found in a purchasable stock.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill aizynthfinder-retrosynthesisInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of aizynthfinder-retrosynthesis
aizynthfinder-retrosynthesis scores 91/100 on our quality scale, 313th of 964 AI & Machine Learning skills we index (top 33%).
Its SKILL.md is 19 KB long, well organised into 40 sections with 16 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so aizynthfinder-retrosynthesis is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
aizynthfinder-retrosynthesis compared with similar skills
All 4 of these similar skills score higher than aizynthfinder-retrosynthesis; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| aizynthfinder-retrosynthesis (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| claude-memby thedotmack | 100 | 96.1k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.3k | 3d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
Frequently asked questions
- How do I install aizynthfinder-retrosynthesis?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill aizynthfinder-retrosynthesis. The install tabs above show the steps for each supported agent. - Which AI agents does aizynthfinder-retrosynthesis work with?
- It is written for Cursor, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is aizynthfinder-retrosynthesis safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is aizynthfinder-retrosynthesis still maintained?
- The repository was last updated 37 days ago, so aizynthfinder-retrosynthesis is actively maintained.
Skill content
View source on GitHubname: "aizynthfinder-retrosynthesis" description: "AiZynthFinder retrosynthetic route planning (CASP) from AstraZeneca Molecular AI. Monte Carlo tree search guided by a template-based neural expansion policy recursively disconnects a target SMILES until precursors are found in a purchasable stock. Covers config.yml (v4 format), aizynthcli batch screening, the AiZynthFinder/AiZynthExpander Python API, one-step disconnections, custom stocks via smiles2stock, scorers, Retro*/breadth-first/DFPN search alternatives, and reading output.json.gz / trees.json. Use for synthesis route planning, synthesizability screening, and building-block/precursor search. For reaction barriers use neb-irc-activation-energy; for 2D reaction scheme drawing use rdkit-chemdraw-cdxml." license: "MIT"
AiZynthFinder Retrosynthesis
Overview
AiZynthFinder performs computer-aided synthesis planning (CASP): a search algorithm — Monte Carlo tree search by default — recursively disconnects a target molecule into precursors, guided by a neural expansion policy that ranks known reaction templates. The search terminates when all precursors are found in a stock (a set of purchasable building blocks) or the maximum depth is reached. Output is a ranked set of reaction trees plus per-target statistics (is_solved, step count, precursors in/out of stock).
Version covered: 4.4.1 (Python 3.10–3.12). The v4 config format differs substantially from v2/v3 as described in the 2020 paper — never copy a config from an old blog post without translating it.
When to Use
- Planning a synthesis route for a designed or purchased target molecule
- Screening a compound library for synthesizability before committing to make-on-demand
- Finding purchasable precursors or building blocks that lead to a scaffold
- Ranking design ideas by route length and by how many precursors fall outside a catalogue
- Enumerating the first retro step only — plausible disconnections without a full tree
- Testing whether a specific bond can be made disconnection-aware (
break_bonds) in a route - Comparing solve rate across two building-block catalogues for the same target set
- Use
torchdruginstead when training a retrosynthesis model rather than running route search - For forward reaction barriers and transition states use
neb-irc-activation-energy; for drawing the resulting scheme userdkit-chemdraw-cdxml
Prerequisites
- Python packages:
aizynthfinder(4.4.x),rdkit,pandas - Data requirements: a stock file (InChIKeys), a trained expansion policy (ONNX model + template CSV), optionally a filter policy
- Environment: Python 3.10–3.12. Default runtime is
onnxruntime; TensorFlow is not needed unless serving remote models or loading legacy.hdf5Keras models.
Check before installing — aizynthcli, download_public_data, and smiles2stock ship with the package and may already be on PATH inside a pixi/conda env. Inside a pixi project, invoke them as pixi run aizynthcli ....
command -v aizynthcli || {
conda create "python>=3.10,<3.13" -n aizynth-env -y
conda activate aizynth-env
python -m pip install "aizynthfinder[all]"
}
[all] adds molbloom (bloom-filter stocks), pymongo, route-distances (route clustering), scipy, and timeout-decorator. Drop it for a lighter install; add [tf] only for TF-serving or .hdf5 models.
Quick Start
from aizynthfinder.aizynthfinder import AiZynthFinder
finder = AiZynthFinder(configfile="config.yml")
finder.stock.select("zinc")
finder.expansion_policy.select("uspto")
finder.target_smiles = "Cc1cccc(c1N(CC(=O)Nc2ccc(cc2)c3ncon3)C(=O)C4CCS(=O)(=O)CC4)C"
finder.tree_search()
finder.build_routes() # required before touching finder.routes
stats = finder.extract_statistics()
print(f"solved={stats['is_solved']} steps={stats['number_of_steps']} "
f"routes={stats['number_of_routes']} time={stats['search_time']:.1f}s")
finder.routes[0]["image"].save("route_top.png")
Workflow
Step 1: Get the Models and Stock
download_public_data fetches the public USPTO models and the ZINC stock subset (several hundred MB, from zenodo.org and figshare.com) and writes a ready-to-use config.yml.
# Skip if the folder already holds the models — this is a large download.
test -f my_folder/config.yml || download_public_data my_folder
ls my_folder
# uspto_model.onnx uspto_templates.csv.gz
# uspto_ringbreaker_model.onnx uspto_ringbreaker_templates.csv.gz
# uspto_filter_model.onnx zinc_stock.hdf5
# config.yml
Step 2: Write or Adjust config.yml
The list short-cut means "template-based strategy, model first, templates second, defaults elsewhere". The same short-cut works for a single filter model path and a single stock file path.
# config.yml — minimal
expansion:
uspto:
- uspto_model.onnx
- uspto_templates.csv.gz
stock:
zinc: zinc_stock.hdf5
# config.yml — explicit form, the settings that matter in practice
search:
algorithm: mcts
algorithm_config:
C: 1.4
use_prior: True
prune_cycles_in_search: True
search_rewards: ["state score"]
max_transforms: 6
iteration_limit: 100
time_limit: 120
return_first: false
exclude_target_from_stock: True
expansion:
uspto:
type: template-based
model: uspto_model.onnx
template: uspto_templates.csv.gz
template_column: retro_template
cutoff_cumulative: 0.995
cutoff_number: 50
use_rdchiral: True
filter:
uspto:
type: quick-filter
model: uspto_filter_model.onnx
filter_cutoff: 0.05
stock:
zinc:
type: inchiset
path: zinc_stock.hdf5
post_processing:
min_routes: 5
max_routes: 25
all_routes: False
Values can be pulled from the environment: iteration_limit: ${ITERATION_LIMIT}.
Step 3: Validate the Target SMILES
An unparseable target burns the whole time limit before failing. Check first.
from rdkit import Chem
smiles = "Cc1cccc(c1N(CC(=O)Nc2ccc(cc2)c3ncon3)C(=O)C4CCS(=O)(=O)CC4)C"
mol = Chem.MolFromSmiles(smiles)
assert mol is not None, f"invalid SMILES: {smiles}"
smiles = Chem.MolToSmiles(mol) # canonicalize
print(f"{smiles} heavy_atoms={mol.GetNumHeavyAtoms()}")
Step 4: Run the Tree Search
select() picks which loaded policies and stocks are active. AiZynthFinder also accepts configdict=<dict> instead of a file — the cleanest way to sweep parameters without writing YAML.
from aizynthfinder.aizynthfinder import AiZynthFinder
finder = AiZynthFinder(configfile="config.yml")
finder.stock.select("zinc")
finder.expansion_policy.select("uspto")
finder.filter_policy.select("uspto") # optional; prunes implausible reactions
finder.target_smiles = smiles
search_time = finder.tree_search()
print(f"search finished in {search_time:.1f}s")
Step 5: Build Routes and Read Statistics
build_routes() extracts reaction trees from the search graph. Nothing in finder.routes exists until it is called.
finder.build_routes()
stats = finder.extract_statistics()
for key in ("is_solved", "number_of_steps", "number_of_routes",
"number_of_precursors", "number_of_precursors_in_stock",
"search_time", "first_solution_time"):
print(f"{key:32s} {stats[key]}")
print("not in stock:", stats["precursors_not_in_stock"])
Step 6: Inspect and Render Routes
finder.routes is a RouteCollection. Show two or three distinct routes, not only the top-scored one.
routes = finder.routes
print(f"{len(routes)} routes, scores: {routes.scores}")
for i in range(min(3, len(routes))):
tree = routes.reaction_trees[i]
leafs = [m.smiles for m in tree.leafs()]
print(f"route {i}: solved={tree.is_solved} "
f"steps={len(list(tree.reactions()))} branched={tree.is_branched()}")
print(f" precursors: {leafs}")
routes.images[i].save(f"route_{i:02d}.png")
routes.jsons[0] # JSON string for the top route
Step 7: Batch Screen with aizynthcli
For hundreds or thousands of targets, use the CLI rather than a Python loop — --nproc splits the input across processes.
# One SMILES per line in smiles.txt
aizynthcli --config config.yml --smiles smiles.txt \
--policy uspto --stocks zinc \
--nproc 8 --checkpoint checkpoint.json.gz \
--output output.json.gz --log_to_file
import pandas as pd
data = pd.read_json("output.json.gz", orient="table")
print(f"solve rate: {data.is_solved.mean():.1%} n={len(data)}")
print(data.loc[data.is_solved, "number_of_steps"].value_counts().sort_index())
print(data.loc[~data.is_solved, ["target", "precursors_not_in_stock"]].head())
Key Parameters
| Parameter | Default | Range / Options | Effect |
|-----------|---------|-----------------|--------|
| search.time_limit | 120 | 30–1800 (s) | Wall-clock budget per target. Raise this first when nothing solves. |
| search.iteration_limit | 100 | 50–1000 | MCTS iterations per target; whichever of time/iterations hits first ends the search. |
| search.max_transforms | 6 | 3–10 | Maximum tree depth (longest route). Deeper searches cost quadratically more. |
| search.return_first | False | True/False | Stop at the first solved route — fast synthesizability yes/no, poor route quality. |
| search.exclude_target_from_stock | True | True/False | Keep True or a purchasable target returns an empty route. |
| search.algorithm_config.C | 1.4 | 0.5–3.0 | UCB exploration/exploitation balance; higher explores more disconnections. |
| search.algorithm_config.search_rewards | ["state score"] | any scorer names | Scorers driving the search; pair with search_rewards_weights for multi-objective. |
| expansion.cutoff_number | 50 | 10–100 | Templates applied per expansion. Widens branching and slows search — tune after the time limit. |
| expansion.cutoff_cumulative | 0.995 | 0.95–0.999 | Cumulative policy probability retained before truncating the template list. |
| expansion.template_column | retro_template | column name | Must match the template file; a mismatch yields silently empty expansions. |
| filter.filter_cutoff | 0.05 | 0.0–0.5 | Feasibility threshold; raising it prunes harder and can make targets unsolvable. |
| post_processing.max_routes | 25 | 5–100 | Routes extracted after the search; all_routes: True returns every solved route. |
Key Concepts
The route score is not a quality score
The state score reflects the fraction of solved precursors and the route length. It was designed to guide the tree search and is largely indiscriminate about whether a route is chemically sensible. Solved routes score near 1.0, unsolved ones typically below 0.8. Never present top_score as a confidence or feasibility measure.
Solve rate is set by the stock and the template library, not the algorithm
The public ZINC subset is far smaller than commercial catalogues; in the original comparison, adding Enamine building blocks found routes for 10 more compounds out of 100. Swapping USPTO for a Reaxys-derived policy changed which compounds solved rather than uniformly improving them. Findability tracks synthetic complexity — an unsolved target means "not found under this stock, this policy, and this budget", not "unsynthesizable".
Reference performance from the paper (100 random ChEMBL compounds, single CPU + single GPU): 55 solved, mean search time 38.7 s, mean time to first solution 7.1 s, mean 2.4 steps and 2.7 precursors.
No conditions are predicted
Reagents, solvents, temperatures, and yields are outside scope. A predicted route is a hypothesis for a chemist to evaluate.
Scorers
Loaded automatically: state score, number of reactions, number of pre-cursors, number of pre-cursors in stock. Also available in aizynthfinder.context.scoring: `average template occur
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
96.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.3kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
