WorldOfTaxonomy
1,000+ open taxonomy systems with 1.2M+ codes and 320K+ crosswalk edges. NAICS, ISIC, NACE, HS, SOC, ISCO, ESCO, CPC, UNSPSC, ICD-10/11, Patent CPC, GDPR, and more - all connected. REST API + MCP server for AI agents. Open source (MIT) by Colaberry Inc.
Install / Use
claude mcp add colaberry -- npx -y github:colaberry/WorldOfTaxonomyIf the server publishes to npm under a different name, use that package instead ā check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of WorldOfTaxonomy
WorldOfTaxonomy scores 84/100 on our quality scale, 564th of 963 AI & Machine Learning skills we index.
Its MCP Server is 16 KB long, well organised into 29 sections with 13 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so WorldOfTaxonomy is actively maintained.
- Our last check on 2026-09-25 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit ā read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk ā read a skill before letting an agent act on it.
WorldOfTaxonomy compared with similar skills
All 4 of these similar skills score higher than WorldOfTaxonomy; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| WorldOfTaxonomy (this skill)by colaberry | 84 | 10 | 2mo ago | MCP Server |
| claude-memby thedotmack | 100 | 95.3k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 89.5k | 18d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.2k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
Frequently asked questions
- How do I install WorldOfTaxonomy?
- Run
claude mcp add colaberry -- npx -y github:colaberry/WorldOfTaxonomy. The install tabs above show the steps for each supported agent. - Which AI agents does WorldOfTaxonomy work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is WorldOfTaxonomy safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is WorldOfTaxonomy still maintained?
- The repository was last updated about 2 months ago, so WorldOfTaxonomy is actively maintained.
Skill content
View source on GitHubWorld Of Taxonomy
<p align="center"> <strong>1,000+ classification systems. 1,305,000+ codes. 326,000+ crosswalk edges.</strong><br> The open-source Rosetta Stone for global industry, trade, occupation, health, and regulatory taxonomies.<br> An open-source project by <a href="https://www.colaberry.ai">Colaberry Inc</a> and <a href="https://www.colaberry.ai">Colaberry Research Labs</a>. </p> <p align="center"> <a href="https://github.com/colaberry/WorldOfTaxonomy/actions/workflows/ci.yml"> <img src="https://github.com/colaberry/WorldOfTaxonomy/actions/workflows/ci.yml/badge.svg" alt="CI" /> </a> <a href="https://opensource.org/licenses/MIT"> <img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="License: MIT" /> [](https://safeskill.dev/scan/colaberry-worldoftaxonomy) </a> <img src="https://img.shields.io/badge/python-3.9%2B-blue.svg" alt="Python 3.9+" /> <img src="https://img.shields.io/badge/next.js-16-black.svg" alt="Next.js 16" /> <img src="https://img.shields.io/badge/MCP-compatible-orange.svg" alt="MCP compatible" /> <img src="https://img.shields.io/badge/systems-1000%2B-purple.svg" alt="1000+ systems" /> <img src="https://img.shields.io/badge/codes-1.3M%2B-green.svg" alt="1.3M+ codes" /> <a href="https://github.com/sponsors/ramdhanyk"> <img src="https://img.shields.io/badge/sponsor-%E2%9D%A4-ea4aaa.svg" alt="Sponsor" /> </a> </p>The Problem
Every country, industry body, and standards organization has its own classification system. When you need to reconcile data across them, you're on your own.
A truck driver in the US is NAICS 484, SOC 53-3032, ISCO-08 8332, NACE 49.4, and ISIC 4923 - five different codes in five different systems that all mean the same thing. Figuring that out manually costs hours. Doing it at scale costs entire teams.
World Of Taxonomy solves this. One queryable graph connects all 1,000+ systems. One API call translates any code to any other system. One MCP server gives AI agents access to the entire taxonomy universe.
Architecture
System Overview
The platform serves four consumer interfaces - a web application, a REST API, an MCP server, and an AI-readable wiki - all backed by a shared PostgreSQL database.
graph TB
subgraph Data["Data Layer"]
PG[(PostgreSQL)]
WIKI["wiki/*.md files"]
end
subgraph Backend["Python Backend"]
INGEST["Ingesters - 1,000+ systems"]
API["FastAPI REST API - /api/v1/*"]
MCP["MCP Server - stdio transport"]
WIKILOADER["Wiki Loader - wiki.py"]
end
subgraph Frontend["Next.js Frontend"]
NEXT["Next.js 15 App Router"]
GUIDE["/guide/* pages"]
end
subgraph Consumers
BROWSER["Web Browsers"]
AIAGENT["AI Agents - Claude, GPT, etc."]
CRAWLER["AI Crawlers - Perplexity, etc."]
DEV["Developer Applications"]
end
INGEST -->|ingest| PG
API -->|query| PG
MCP -->|query| PG
WIKILOADER -->|read| WIKI
MCP -->|instructions| WIKILOADER
NEXT -->|proxy /api/*| API
NEXT -->|read| WIKI
GUIDE -->|render| WIKI
BROWSER --> NEXT
BROWSER --> GUIDE
AIAGENT --> MCP
CRAWLER -->|/llms-full.txt| NEXT
DEV --> API
Ingestion Pipeline
Each of the 1,000+ systems has a dedicated ingester that fetches from authoritative sources and loads into three core tables.
graph TD
subgraph Sources["Official Sources"]
CSV["CSV files - NAICS, ISIC"]
XLSX["Excel files - NACE, ANZSIC"]
HTML["HTML/PDF - SIC, NIC"]
CURATED["Expert-Curated - Domain taxonomies"]
end
subgraph Pipeline["Ingestion Pipeline"]
PARSE["Parse and Validate"]
UPSERT["Upsert Nodes into classification_node"]
XWALK["Build Crosswalks into equivalence"]
PROV["Set Provenance - 4-tier audit"]
end
subgraph DB["Database Tables"]
SYS["classification_system - 1,000+ systems"]
NODE["classification_node - 1.3M+ nodes"]
EQUIV["equivalence - 326K+ edges"]
end
CSV --> PARSE
XLSX --> PARSE
HTML --> PARSE
CURATED --> PARSE
PARSE --> UPSERT
PARSE --> XWALK
PARSE --> PROV
UPSERT --> NODE
XWALK --> EQUIV
PROV --> SYS
SYS --- NODE
NODE --- EQUIV
API Request Flow
sequenceDiagram
participant C as Client
participant RL as Rate Limiter
participant AUTH as Auth Layer
participant R as Router
participant Q as Query Layer
participant DB as PostgreSQL
C->>RL: GET /api/v1/search?q=physician
RL->>RL: Check rate - 30/min anon, 1000/min auth
RL->>AUTH: Forward request
AUTH->>AUTH: Validate JWT or API key
AUTH->>R: Authenticated request
R->>Q: search(conn, query, limit)
Q->>DB: SELECT with ts_vector query
DB-->>Q: Matching nodes
Q-->>R: Results with system context
R-->>C: JSON response
MCP Session Lifecycle
sequenceDiagram
participant AI as AI Agent
participant MCP as MCP Server
participant WIKI as Wiki Loader
participant DB as PostgreSQL
AI->>MCP: initialize - JSON-RPC
MCP->>WIKI: build_wiki_context()
WIKI-->>MCP: Structural knowledge - ~15K tokens
MCP-->>AI: serverInfo + instructions + capabilities
Note over AI: Agent now knows all 1,000+ systems and crosswalk topology
AI->>MCP: tools/call search_classifications
MCP->>DB: Query nodes
DB-->>MCP: Results
MCP-->>AI: Tool response as JSON
AI->>MCP: resources/read taxonomy://wiki/crosswalk-map
MCP->>WIKI: load_wiki_page - crosswalk-map
WIKI-->>MCP: Full markdown content
MCP-->>AI: Resource content
Directory Structure
World Of Taxonomy/
āāā world_of_taxonomy/
ā āāā api/ # FastAPI REST API (lifespan pool, rate limiting)
ā ā āāā routers/ # systems, nodes, search, equivalences, countries, auth, crosswalk_graph
ā ā āāā schemas.py # Pydantic response models
ā āāā mcp/ # MCP server (stdio transport, 26 tools)
ā āāā ingest/ # One ingester per system (100+ files)
ā ā āāā naics.py # Downloads from Census Bureau
ā ā āāā nace_derived.py # EU national adaptations (copy NACE + equivalences)
ā ā āāā isic_derived.py # LATAM/Asia/Africa adaptations
ā ā āāā crosswalk_*.py # 20+ crosswalk ingesters
ā āāā query/ # Query layer (browse, search, equivalence)
ā āāā schema.sql # Core tables
ā āāā schema_auth.sql # Auth tables
āāā frontend/ # Next.js 15 + TypeScript + Tailwind + shadcn/ui
ā āāā src/app/ # Home, Explore, System, Dashboard, Crosswalk Explorer, Guide
āāā wiki/ # Curated guides (serves web, MCP, llms.txt, and API)
āāā tests/ # pytest (test_wot schema isolation, never touches production)
āāā data/ # Downloaded source files (gitignored, re-downloadable)
Database (PostgreSQL):
classification_system -- 1,000 rows: id, name, region, authority, node_count
classification_node -- 1.2M rows: system_id, code, title, level, parent_code
equivalence -- 326K rows: source_system, source_code, target_system, target_code, match_type (edge_kind is computed on read)
country_system_link -- 27K rows: country_code, system_id, relevance ('official'|'regional'|'recommended')
Quick Start
Docker (recommended - runs in under 2 minutes):
git clone https://github.com/colaberry/WorldOfTaxonomy.git
cd World Of Taxonomy
docker compose up
Open http://localhost:3000. The API is at http://localhost:8000.
Then ingest your first systems:
# Core global systems (~3 minutes)
docker compose exec backend python3 -m world_of_taxonomy ingest naics
docker compose exec backend python3 -m world_of_taxonomy ingest isic
docker compose exec backend python3 -m world_of_taxonomy ingest crosswalk
# Everything (~30-45 minutes)
docker compose exec backend python3 -m world_of_taxonomy ingest all
Python only (bring your own PostgreSQL):
pip install -e .
cp .env.example .env # set DATABASE_URL and JWT_SECRET
python3 -m world_of_taxonomy init
python3 -m world_of_taxonomy ingest naics
python3 -m uvicorn world_of_taxonomy.api.app:create_app --factory --port 8000
API in 60 Seconds
# Translate NAICS 4841 (general freight trucking) to all equivalent systems
curl "http://localhost:8000/api/v1/systems/naics_2022/nodes/4841/translations"
# Search for "hospital" across all 1,000+ systems simultaneously
curl "http://localhost:8000/api/v1/search?q=hospital&grouped=true"
# Get every classification system applicable to Germany
curl "http://localhost:8000/api/v1/countries/DE"
# Find all codes in NACE with no mapping to NAICS
curl "http://localhost:8000/api/v1/diff?a=nace_rev2&b=naics_2022"
# Full text search within a specific system
curl "http://localhost:8000/api/v1/search?q=logistics&system=isco_08"
Python client:
import httplib2, json
base = "http://localhost:8000/api/v1"
# Get all systems for France
r = httplib2.Http().request(f"{base}/countries/FR")
profile = json.loads(r[1])
# {'official': 'naf_rev2', 'regional': 'nace_rev2', 'recommended': ['isic_rev4', ...]}
# Translate a code
r = httplib2.Http().request(f"{base}/systems/naics_2022/nodes/5415/translations")
translations = json.loads(r[1])
# {'nace_rev2': '62.01', 'isic_rev4': '6201', 'sic_1987': '7371', ...}
Use With Claude / AI Agents (MCP)
World Of Taxonomy ships with a Model Context Protocol server. Add it to Claude Desktop and your AI gets instant access to all 1,000+ systems as structured tools.
~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"world-of-taxonomy": {
"command": "python3",
"args": ["-m", "world_of_taxonomy", "mcp"],
"env": {
"DATABASE_URL": "your-database-url"
}
}
}
}
26 tools available, including:
| Tool | What it does |
|------|-------------|
| translate_code | Translate any code to a target system |
| translate_across_all_systems | One code -> all 1,000+ systems at once |
| search_classifications | Full-text search across all codes |
| get_country_taxonomy_profile | Official + recommended systems for any country |
| compare_sector | Side-by-side root nodes across two systems |
| get_system_diff | Codes in system A with no mapping to B |
| explore_industry_tree | Browse hierarchy with context |
| find_by_keyword_all_systems | Search grouped by system |
Example prompt to Claude:
"I have a dataset with NACE codes. Convert every unique code to NAICS and ISIC equivalents and flag any that have no crosswalk."
What's Covered
16 categories. 1,000+ systems. Every major region.
| Category | Systems | Highlights |
|----------|---------|-----------|
| Industry | 68 | NAICS, ISIC, NACE + 58 national adaptations (EU, LATAM, Asia, Africa) |
| Life Sciences | 108 | ICD-11, ICD-10-CM/PCS, LOINC (102K), MeSH, SNOMED, NDC, NCI Thesaurus (211K) |
| Domain Deep-Dives | 434 | Plain-language sector vocabularies for 40+ verticals, all bridged to NAICS / ISIC / NACE via sector anchors (derived:sector_anchor:v1) |
| Regulatory | 80+ | GDPR, FDA, SOX, HIPAA, ISO standards, EU directives, NIST frameworks |
| Occupational | 10 | SOC, ISCO-08, ESCO (14K skills), O*NET, ANZSCO, NOC, KldB, ROME |
| Product / Trade | 11 | HS 2022, UNSPSC (77K codes), CPC, SITC, HTS, Schedule B, ECCN |
| Research & Knowledge | 8 | FORD, JEL, LCC, PACS, MSC, ACM CCS, arXiv, ANZSRC |
| Financial / Investment | 7 | GICS, ICB, CFI (ISO 10962), COFOG, COICOP, GHG Protocol, Patent CPC |
| Geographic | 7 | ISO 3166-1/2, UN M.49, EU NUTS, US FIPS, World Bank income groups |
| Education | 3 | ISCED 2011, ISCED-F 2013, CIP 2020 |
249 countries are profiled with their official, regional, and recommended systems.
Use Cases
- Data engineering: Reconcile supplier data (NAICS) with EU reporting (
Truncated for display ā read the full file on GitHub.
Related Skills
claude-mem
95.3kPersistent Context Across Sessions for Every Agent ā Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
89.5kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu ā one CLI, zero API fees.
Understand-Anything
85.2kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit ā see the Safety scan above for what the skill file itself contains.
