arboreto-grn-inference
GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inferenceInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of arboreto-grn-inference
arboreto-grn-inference scores 91/100 on our quality scale, 1163rd of 4,619 Development & Engineering skills we index (top 26%).
Its SKILL.md is 21 KB long, well organised into 51 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so arboreto-grn-inference is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
arboreto-grn-inference compared with similar skills
All 4 of these similar skills score higher than arboreto-grn-inference; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| arboreto-grn-inference (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.0k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install arboreto-grn-inference?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference. The install tabs above show the steps for each supported agent. - Which AI agents does arboreto-grn-inference work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is arboreto-grn-inference safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is arboreto-grn-inference still maintained?
- The repository was last updated 37 days ago, so arboreto-grn-inference is actively maintained.
Skill content
View source on GitHubname: "arboreto-grn-inference" description: "GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component." license: "BSD-3-Clause"
Arboreto GRN Inference
Overview
Arboreto infers gene regulatory networks (GRNs) from gene expression data using parallelized tree-based regression. For each target gene, it trains a regression model with all other genes (or a specified TF list) as features and emits TF-target-importance triplets. It provides two interchangeable algorithms -- GRNBoost2 (gradient boosting, fast) and GENIE3 (Random Forest, classic) -- sharing identical input/output formats. Computation is Dask-parallelized, scaling from laptop cores to HPC clusters.
When to Use
- Inferring transcription factor-to-target gene regulatory relationships from bulk RNA-seq expression data
- Building gene regulatory networks from single-cell RNA-seq count matrices (cells as rows, genes as columns)
- Generating the adjacency matrix (Step 1) of the pySCENIC regulatory analysis pipeline
- Comparing regulatory network structure across experimental conditions (e.g., control vs treatment)
- Producing consensus regulatory networks by running inference across multiple random seeds
- Validating GRN results by comparing GRNBoost2 and GENIE3 outputs on the same dataset
- For downstream regulon identification and activity scoring, use arboreto output with pySCENIC
- For single-cell preprocessing (QC, normalization, clustering) before GRN inference, use scanpy-scrna-seq
Prerequisites
- Python packages:
arboreto,pandas,numpy,dask,distributed,scikit-learn,scipy - Data requirements: Gene expression matrix (genes as columns, observations as rows) in TSV/CSV; optionally a TF list file (one gene name per line)
- Environment: Python 3.8+; optional
networkx,matplotlibfor visualization
pip install arboreto distributed networkx matplotlib
Pre-flight Interview
Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: algorithm
kind: required
source: user
ask: "Infer the network with gradient boosting (faster) or random forests (the original GENIE3 formulation)?"
default: "GRNBoost2"
- id: D2
param: expressionMatrix
kind: required
source: upstream
ask: "Which processed expression matrix, and which cells or samples within it, should the network be inferred from?"
default: null
- id: D3
param: regulatorList
kind: required
source: literature
ask: "Restrict candidate regulators to a curated transcription-factor list, or let every gene be a candidate regulator?"
default: "a species-matched TF list"
- id: D4
param: edgeThreshold
kind: required
source: user
ask: "The output ranks every regulator-target pair - how many of the top edges should be kept as the network?"
default: null
- id: D5
param: seed
kind: never_ask
source: data
reason: "Fixes the draw for reproducibility; does not change what the data supports"
default: 0
- id: D6
param: daskClient
kind: never_ask
source: data
reason: "Execution backend; affects runtime only"
default: "local"
D3 is the decision that separates a usable network from a hairball. Leaving regulators unrestricted lets any correlated gene appear as a regulator, and the result still returns a fully populated ranked table. D4 has no safe default: the algorithm scores all pairs, so where the list is cut is the user's call about the network they intend to interpret.
Quick Start
Complete GRN inference in a single block. The if __name__ == '__main__': guard is required because Dask spawns worker processes via multiprocessing.
import pandas as pd
from arboreto.algo import grnboost2
from arboreto.utils import load_tf_names
if __name__ == '__main__':
# Load expression matrix (observations x genes)
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
tf_names = load_tf_names('tf_list.txt') # optional TF filter
# Infer GRN (uses all local cores by default)
network = grnboost2(expression_data=expression_matrix,
tf_names=tf_names, seed=777)
# Filter top links and save
top_network = network[network['importance'] > 1.0]
top_network.to_csv('grn_output.tsv', sep='\t', index=False, header=False)
print(f"Inferred {len(network)} links, kept {len(top_network)} above threshold")
# Example: Inferred 185432 links, kept 12876 above threshold
Workflow
Step 1: Load Expression Data
Arboreto accepts a pandas DataFrame (recommended) or NumPy array. Rows are observations (cells or samples), columns are genes. Gene names must be column headers.
import pandas as pd
# From TSV (genes as columns, observations as rows)
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
print(f"Shape: {expression_matrix.shape} "
f"({expression_matrix.shape[0]} observations x {expression_matrix.shape[1]} genes)")
# Example: Shape: (5000, 18654) (5000 observations x 18654 genes)
# From AnnData (e.g., after scanpy preprocessing)
import anndata as ad
adata = ad.read_h5ad('preprocessed.h5ad')
expression_matrix = pd.DataFrame(
adata.X.toarray() if hasattr(adata.X, 'toarray') else adata.X,
columns=adata.var_names.tolist()
)
print(f"Converted AnnData: {expression_matrix.shape}")
# Example: Converted AnnData: (5000, 18654)
Step 2: Load Transcription Factor List
Providing a TF list restricts regulators to known transcription factors, reducing computation time and improving biological relevance. If omitted, all genes are treated as potential regulators.
from arboreto.utils import load_tf_names
# From file (one TF name per line)
tf_names = load_tf_names('human_tfs.txt')
print(f"Loaded {len(tf_names)} transcription factors")
# Example: Loaded 1639 transcription factors
# Or define directly
tf_names = ['MYC', 'TP53', 'SOX2', 'NANOG', 'POU5F1']
# Verify TFs exist in expression matrix columns
tf_in_data = [tf for tf in tf_names if tf in expression_matrix.columns]
print(f"TFs found in expression data: {len(tf_in_data)}/{len(tf_names)}")
# Example: TFs found in expression data: 1583/1639
Step 3: Configure Dask Client (Optional)
By default arboreto creates an internal Dask client using all local cores. Create an explicit client for resource control, monitoring, or cluster deployment.
from distributed import LocalCluster, Client
# Custom local client with resource limits
local_cluster = LocalCluster(
n_workers=8,
threads_per_worker=1, # avoid GIL contention in scikit-learn
memory_limit='4GB' # per worker
)
client = Client(local_cluster)
print(f"Dashboard: {client.dashboard_link}")
# Example: Dashboard: http://127.0.0.1:8787/status
Step 4: Run GRN Inference
Call grnboost2() (recommended) or genie3(). Both share the same signature and output format.
from arboreto.algo import grnboost2
if __name__ == '__main__':
network = grnboost2(
expression_data=expression_matrix,
tf_names=tf_names, # 'all' if no TF list
client_or_address=client, # omit to use default local scheduler
seed=777, # for reproducibility
verbose=True # print progress
)
print(f"Inferred {len(network)} regulatory links")
print(network.head())
# Example:
# TF target importance
# 0 MYC CDK4 3.214
# 1 MYC CCND1 2.871
# 2 TP53 CDKN1A 2.654
Step 5: Filter Results
Raw output contains links for every TF-target pair with non-zero importance. Filter to retain high-confidence regulatory interactions.
# Strategy 1: Importance threshold
threshold = 1.0
filtered = network[network['importance'] > threshold]
print(f"Threshold {threshold}: {len(filtered)} links "
f"({len(filtered)/len(network)*100:.1f}% retained)")
# Example: Threshold 1.0: 12876 links (6.9% retained)
# Strategy 2: Top N links per target gene
top_n = 10
top_per_target = (network.groupby('target')
.apply(lambda g: g.nlargest(top_n, 'importance'))
.reset_index(drop=True))
print(f"Top {top_n} per target: {len(top_per_target)} links")
# Example: Top 10 per target: 186540 links
Step 6: Save Network
Save as a TSV file with three columns: TF, target, importance.
# Save full network
network.to_csv('full_network.tsv', sep='\t', index=False)
print(f"Saved full network: {len(network)} links")
# Save filtered network (without header for pySCENIC compatibility)
filtered.to_csv('filtered_network.tsv', sep='\t', index=False, header=False)
print(f"Saved filtered network: {len(filtered)} links")
# Clean up Dask client if explicitly created
client.close()
local_cluster.close()
Step 7: Visualize Results (Optional)
Plot the top regulatory interactions as a directed network graph.
import networkx as nx
import matplotlib.pyplot as plt
# Build directed graph from top links
top_links = network.nlargest(50, 'importance')
G = nx.from_pandas_edgelist(
top_links, source='TF', target='target',
edge_attr='importance', create_using=nx.DiGraph()
)
# Draw network
fig, ax = plt.subplots(figsize=(12, 10))
pos = nx.spring_layout(G, k=2, seed=42)
nx.draw_networkx_nodes(G, pos, node_size=300, node_color='lightblue', ax=ax)
nx.draw_networkx_labels(G, pos, font_size=7, ax=ax)
edges = nx.draw_networkx_edges(
G, pos, edge_color=[G[u][v]['importance'] for u, v in G.edges()],
edge_cmap=plt.cm.Reds, width=1.5, arrows=True, ax=ax
)
plt.colorbar(edges, ax=ax, label='Importance')
ax.set_title('Top 50 Regulatory Interactions')
plt.tight_layout()
plt.savefig('grn_network.png', dpi=300, bbox_inches='tight')
print("Saved grn_network.png")
Key Parameters
| Parameter | Default | Range / Options | Effect |
|-----------|---------|-----------------|--------|
| expression_data | (required) | DataFrame or ndarray | Expression matrix, observations x genes |
| tf_names | 'all' | list of str or 'all' | Restrict regulators to known TFs; 'all' uses every gene |
| gene_names | None | list of str | Required when expression_data is a NumPy array |
| client_or_address | 'local' | Client, str, or 'local' | Dask client instance or scheduler address |
| seed | None | int | Random seed for reproducibility; always set for replicable results |
| verbose | False | bool | Print progress messages during inference |
Key Concepts
Algorithm Selection Guide
Both algorithms follow the same multiple-regression strategy (train one model per target gene, extract feature importances) but differ in the underlying regressor.
| Feature | GRNBoost2 | GENIE3 |
|---------|-----------|--------|
| Method | Stochastic gradient boosting with early stopping | Random Forest (or ExtraTrees) |
| Speed | Fast -- optimized for large datasets | Slower -- higher per-model cost |
| Memory | Lower (early stopping limits tree depth) | Higher (full forests per target) |
| Best for | Default choice; 10k+ observations, single-cell scale | Comparison with published GENIE3 results; validation |
| Import | from arboreto.algo import grnboost2 | from arboreto.algo import genie3 |
| Output format | TF, target, importance (identical) | TF, target, importance (identical) |
Decision rule: Start with GRNBoost2. Use GENIE3 only when reproducing published GENIE3 analyses or as an independent validation of GRNBoost2 results.
Output Format
The network DataFrame has three columns, sorted by descending importance:
| Column | Type | Description |
|--------|------|-------------|
| TF | str | Transcription factor (regulator) gene name |
| target
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
45.0kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
