interpro-database
Query InterPro REST API for protein domain architecture, family classification, and member-DB integration. Search entries, retrieve a protein's domains, list family members, get taxonomic distribution, link to PDB. Unifies Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill interpro-databaseInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of interpro-database
interpro-database scores 91/100 on our quality scale, 204th of 573 Data & Analytics skills we index (top 36%).
Its SKILL.md is 30 KB long, well organised into 63 sections with 18 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so interpro-database is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
interpro-database compared with similar skills
All 4 of these similar skills score higher than interpro-database; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| interpro-database (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install interpro-database?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill interpro-database. The install tabs above show the steps for each supported agent. - Which AI agents does interpro-database work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is interpro-database safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is interpro-database still maintained?
- The repository was last updated 37 days ago, so interpro-database is actively maintained.
Skill content
View source on GitHubname: "interpro-database" description: "Query InterPro REST API for protein domain architecture, family classification, and member-DB integration. Search entries, retrieve a protein's domains, list family members, get taxonomic distribution, link to PDB. Unifies Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam. Use uniprot-protein-database for sequences; pdb-database for 3D structures." license: "CC-BY-4.0"
InterPro Database
Overview
InterPro is the EBI's integrated protein family, domain, and functional site database. It consolidates signatures from 13 member databases (Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam, and others) into unified InterPro entries, each describing a homologous superfamily, domain, family, repeat, or conserved site. The REST API at https://www.ebi.ac.uk/interpro/api/ is free and requires no authentication.
When to Use
- Identifying all domains and families present in a protein by UniProt accession (domain architecture)
- Searching for proteins that contain a specific domain or belong to a specific family
- Finding the taxonomic distribution of organisms that encode a given domain or family
- Cross-linking a domain to experimental 3D structures in the PDB
- Checking which source databases (Pfam, PANTHER, SMART, etc.) cover an InterPro entry
- Discovering InterPro entries by keyword (e.g., "kinase domain") when you do not yet know the accession
- For protein sequence retrieval, functional annotations (GO, pathways, active sites), and ID mapping use
uniprot-protein-database - For downloading domain-aligned sequences or building HMM profiles use
Pfamdirectly; InterPro is the meta-layer
Prerequisites
- Python packages:
requests,pandas,matplotlib - Data requirements: UniProt accessions (e.g.,
P04637) or InterPro accessions (e.g.,IPR011009) - Environment: internet connection; no API key required
- Rate limits: no published hard limit; use
time.sleep(1.0)between requests for batch queries; paginate with?cursor=or?page_size=
pip install requests pandas matplotlib
Quick Start
import requests
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def interpro_get(path: str, params: dict = None) -> dict:
"""Send a GET request to the InterPro API and return parsed JSON."""
r = requests.get(
f"{INTERPRO_BASE}/{path}",
params=params,
headers={"Accept": "application/json"},
timeout=30
)
r.raise_for_status()
return r.json()
# Get domain architecture for TP53 (P04637)
# Note: `protein/uniprot/{acc}/` returns only {metadata}; the entries-per-protein
# data lives at `entry/interpro/protein/uniprot/{acc}/` and is keyed `results`.
data = interpro_get("entry/interpro/protein/uniprot/P04637/")
entries = data.get("results", [])
print(f"InterPro entries for TP53: {data.get('count')} (this page: {len(entries)})")
for e in entries[:4]:
m = e["metadata"]
print(f" {m['accession']} {m['type']:<25} {m['name']}")
# InterPro entries for TP53: 9
# IPR002117 family p53 tumour suppressor family
# IPR036674 homologous_superfamily p53-like tetramerisation domain superfamily
Core API
Query 1: Entry Search
Search for InterPro entries by name keyword or fetch a specific entry by accession.
import requests
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def search_entries(query: str, entry_type: str = None,
page_size: int = 20) -> list:
"""Search InterPro entries by keyword; optionally filter by type."""
params = {"search": query, "page_size": page_size}
if entry_type:
params["type"] = entry_type # family, domain, homologous_superfamily, repeat, site
r = requests.get(
f"{INTERPRO_BASE}/entry/interpro/",
params=params,
headers={"Accept": "application/json"},
timeout=30
)
r.raise_for_status()
return r.json().get("results", [])
hits = search_entries("serine kinase", entry_type="domain")
print(f"InterPro domain entries matching 'serine kinase': {len(hits)}")
for h in hits[:5]:
m = h["metadata"]
print(f" {m['accession']} {m['type']:<10} {m['name']}")
# InterPro domain entries matching 'serine kinase': 8
# IPR000719 domain Protein kinase domain
# IPR008271 domain Serine/threonine/tyrosine kinase, active site
# Fetch a specific InterPro entry by accession
r = requests.get(
f"{INTERPRO_BASE}/entry/interpro/IPR000719/",
headers={"Accept": "application/json"},
timeout=30
)
r.raise_for_status()
meta = r.json()["metadata"]
print(f"Accession : {meta['accession']}")
print(f"Name : {meta['name']}")
print(f"Type : {meta['type']}")
print(f"Member DBs : {list(meta.get('member_databases', {}).keys())}")
go_terms = meta.get("go_terms", [])
print(f"GO terms : {[g['identifier'] for g in go_terms[:3]]}")
# Accession : IPR000719
# Name : Protein kinase domain
# Type : domain
# Member DBs : ['pfam', 'smart', 'cdd', 'ncbifam', 'panther']
# GO terms : ['GO:0004672', 'GO:0005524', 'GO:0006468']
Query 2: Protein Domain Architecture
Retrieve all InterPro entries (domains, families, sites) matched in a protein by UniProt accession.
import requests
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def get_protein_domain_architecture(uniprot_acc: str) -> dict:
"""Return all InterPro entry matches for a protein. Uses the
`entry/interpro/protein/uniprot/{acc}/` endpoint, which returns
{count, next, previous, results}. Each result has metadata + a
nested `proteins[0].entry_protein_locations` for the per-protein match."""
r = requests.get(
f"{INTERPRO_BASE}/entry/interpro/protein/uniprot/{uniprot_acc}/",
headers={"Accept": "application/json"},
timeout=60
)
r.raise_for_status()
return r.json()
data = get_protein_domain_architecture("P04637") # TP53
results = data.get("results", [])
# Pull length/source from the first match's nested protein record
prot0 = results[0]["proteins"][0] if results and results[0].get("proteins") else {}
print(f"Protein length : {prot0.get('protein_length')}")
print(f"Source DB : {prot0.get('source_database')}")
print(f"InterPro entries: {data.get('count')}")
for entry in results[:6]:
m = entry["metadata"]
# Locations are nested under proteins[0].entry_protein_locations
locs = entry["proteins"][0].get("entry_protein_locations", []) if entry.get("proteins") else []
loc_str = ", ".join(
f"{frag['start']}-{frag['end']}"
for loc in locs for frag in loc.get("fragments", [])
)
print(f" {m['accession']} {m['type']:<25} {m['name'][:35]:<35} [{loc_str}]")
# Compare domain architectures of two proteins side-by-side
import pandas as pd
def domain_set(uniprot_acc: str) -> set:
data = get_protein_domain_architecture(uniprot_acc)
return {e["metadata"]["accession"] for e in data.get("results", [])}
brca1_domains = domain_set("P38398") # BRCA1
tp53_domains = domain_set("P04637") # TP53
shared = brca1_domains & tp53_domains
unique_brca1 = brca1_domains - tp53_domains
unique_tp53 = tp53_domains - brca1_domains
print(f"Shared InterPro entries: {len(shared)}")
print(f"BRCA1-unique : {len(unique_brca1)}")
print(f"TP53-unique : {len(unique_tp53)}")
Query 3: Entry Proteins
List proteins that contain a specific InterPro entry (family or domain).
import requests, time
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def get_entry_proteins(interpro_acc: str, reviewed_only: bool = True,
page_size: int = 50) -> list:
"""Return proteins (UniProt) containing a given InterPro entry.
Path order is `/protein/{db}/entry/interpro/{IPR}/` — the inverse
`entry/interpro/{IPR}/protein/{db}/` times out (408) on large families."""
db = "reviewed" if reviewed_only else "uniprot"
r = requests.get(
f"{INTERPRO_BASE}/protein/{db}/entry/interpro/{interpro_acc}/",
params={"page_size": page_size},
headers={"Accept": "application/json"},
timeout=60
)
r.raise_for_status()
return r.json().get("results", [])
proteins = get_entry_proteins("IPR011009") # Protein kinase-like domain SF
print(f"Reviewed proteins with IPR011009 (page 1): {len(proteins)}")
for p in proteins[:4]:
m = p["metadata"]
# metadata fields: accession, gene, length, name, source_database, source_organism
print(f" {m['accession']} {(m.get('gene') or ''):<8} "
f"len={m.get('length', '?')} "
f"org={(m.get('source_organism') or {}).get('scientificName', '')[:30]}")
# Paginate all proteins for a family using cursor
def get_all_entry_proteins(interpro_acc: str,
reviewed_only: bool = True) -> list:
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
db = "reviewed" if reviewed_only else "uniprot"
# Path-inverted: /protein/{db}/entry/interpro/{IPR}/ is the working order
url = f"{INTERPRO_BASE}/protein/{db}/entry/interpro/{interpro_acc}/"
all_proteins = []
params = {"page_size": 200}
while url:
r = requests.get(url, params=params,
headers={"Accept": "application/json"}, timeout=60)
r.raise_for_status()
data = r.json()
all_proteins.extend(data.get("results", []))
url = data.get("next")
params = None # next URL already has params encoded
if url:
time.sleep(1.0)
return all_proteins
proteins = get_all_entry_proteins("IPR000719") # Protein kinase domain
print(f"Total reviewed proteins with protein kinase domain: {len(proteins)}")
Query 4: Entry Taxonomy
Get the taxonomic distribution of proteins annotated with a given InterPro entry.
import requests
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def get_entry_taxonomy(interpro_acc: str,
page_size: int = 50) -> list:
"""Return taxonomic summary for proteins in a given InterPro entry.
Path-inverted: `/taxonomy/uniprot/entry/interpro/{IPR}/`. Each result
has `metadata` (taxon: accession=taxId, name, parent, children, rank)
and `entries[]` (representative protein-match locations for that taxon)."""
r = requests.get(
f"{INTERPRO_BASE}/taxonomy/uniprot/entry/interpro/{interpro_acc}/",
params={"page_size": page_size},
headers={"Accept": "application/json"},
timeout=90
)
r.raise_for_status()
return r.json().get("results", [])
# Use a smaller entry (p53 DBD); IPR000719 (kinase) has ~270k taxa and times out.
taxa = get_entry_taxonomy("IPR011615")
print(f"Top taxa for IPR011615 (p53 DNA-binding domain):")
for t in taxa[:8]:
m = t["metadata"]
print(f" taxId={m['accession']:>10} {m.get('name', ''):<30} "
f"rank={m.get('rank') or 'n/a'}")
Query 5: Structure Integration
Retrieve PDB structures associated with an InterPro entry.
import requests
INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
def get_entry_structures(interpro_acc: str, page_size: int = 25) -> list:
"""Return PDB structures that include a match to a given InterPro entry.
Path-inverted: `/structure/pdb/entry/interpro/{IPR}/`. The flat form with
`?entry_interpro=...` is silently slow / 408s on this resource."""
r = requests.get(
f"{INTERPRO_BASE}/structure/pdb/entry/interpro/{interpro_acc}/",
params={"page_size": page_size},
headers={"Accept": "application/json"},
timeout=60
)
r.raise_for_status()
return r.json().get("results", [])
structures = get_entry_structures("IPR011009") # Protein kinase-like SF
print(f"PDB structures linked to IPR011009 (page 1): {len(structures)}")
for s in structures[:5]:
m = s["metadata"]
print(f" {m['
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
