hierarchical-taxonomy-clustering
Build unified multi-level category taxonomy from hierarchical product category paths from any e-commerce companies using embedding-based recursive clustering with intelligent category naming via weighted word frequency analysis.
Install / Use
npx skills add benchflow-ai/skillsbench --skill hierarchical-taxonomy-clusteringInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of hierarchical-taxonomy-clustering
hierarchical-taxonomy-clustering scores 86/100 on our quality scale, 1428th of 3,055 Automation skills we index (top 47%).
Its SKILL.md is 3.8 KB long, well organised into 11 sections with 1 code example: a solid amount of guidance for an agent.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so hierarchical-taxonomy-clustering is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
hierarchical-taxonomy-clustering compared with similar skills
All 4 of these similar skills score higher than hierarchical-taxonomy-clustering; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| hierarchical-taxonomy-clustering (this skill)by benchflow-ai | 86 | 1.8k | 2mo ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 87.6k | 16d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.7k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.1k | 1d ago | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 9d ago | SKILL.md |
Frequently asked questions
- How do I install hierarchical-taxonomy-clustering?
- Run
npx skills add benchflow-ai/skillsbench --skill hierarchical-taxonomy-clustering. The install tabs above show the steps for each supported agent. - Which AI agents does hierarchical-taxonomy-clustering work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is hierarchical-taxonomy-clustering safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is hierarchical-taxonomy-clustering still maintained?
- The repository was last updated about 2 months ago, so hierarchical-taxonomy-clustering is actively maintained.
Skill content
View source on GitHubname: hierarchical-taxonomy-clustering description: Build unified multi-level category taxonomy from hierarchical product category paths from any e-commerce companies using embedding-based recursive clustering with intelligent category naming via weighted word frequency analysis.
Hierarchical Taxonomy Clustering
Create a unified multi-level taxonomy from hierarchical category paths by clustering similar paths and automatically generating meaningful category names.
Problem
Given category paths from multiple sources (e.g., "electronics -> computers -> laptops"), create a unified taxonomy that groups similar paths across sources, generates meaningful category names, and produces a clean N-level hierarchy (typically 5 levels). The unified category taxonomy could be used to do analysis or metric tracking on products from different platform.
Methodology
- Hierarchical Weighting: Convert paths to embeddings with exponentially decaying weights (Level i gets weight 0.6^(i-1)) to signify the importance of category granularity
- Recursive Clustering: Hierarchically cluster at each level (10-20 clusters at L1, 3-20 at L2-L5) using cosine distance
- Intelligent Naming: Generate category names via weighted word frequency + lemmatization + bundle word logic
- Quality Control: Exclude all ancestor words (parent, grandparent, etc.), avoid ancestor path duplicates, clean special characters
Output
DataFrame with added columns:
unified_level_1: Top-level category (e.g., "electronic | device")unified_level_2: Second-level category (e.g., "computer | laptop")unified_level_3throughunified_level_N: Deeper levels
Category names use | separator, max 5 words, covering 70%+ of records in each cluster.
Installation
pip install pandas numpy scipy sentence-transformers nltk tqdm
python -c "import nltk; nltk.download('wordnet'); nltk.download('omw-1.4')"
4-Step Pipeline
Step 1: Load, Standardize, Filter and Merge (step1_preprocessing_and_merge.py)
- Input: List of (DataFrame, source_name) tuples, each of the with
category_pathcolumn - Process: Per-source deduplication, text cleaning (remove &/,/'/-/quotes,'and' or "&", "," and so on, lemmatize words as nouns), normalize delimiter to
>, depth filtering, prefix removal, then merge all sources. source_level should reflect the processed version of the source level name - Output: Merged DataFrame with
category_path,source,depth,source_level_1throughsource_level_N
Step 2: Weighted Embeddings (step2_weighted_embedding_generation.py)
- Input: DataFrame from Step 1
- Output: Numpy embedding matrix (n_records × 384)
- Weights: L1=1.0, L2=0.6, L3=0.36, L4=0.216, L5=0.1296 (exponential decay 0.6^(n-1))
- Performance: For ~10,000 records, expect 2-5 minutes. Progress bar will show encoding status.
Step 3: Recursive Clustering (step3_recursive_clustering_naming.py)
- Input: DataFrame + embeddings from Step 2
- Output: Assignments dict {index → {level_1: ..., level_5: ...}}
- Average linkage + cosine distance, 10-20 clusters at L1, 3-20 at L2-L5
- Word-based naming: weighted frequency + lemmatization + coverage ≥70%
- Performance: For ~10,000 records, expect 1-3 minutes for hierarchical clustering and naming. Be patient - the system is working through recursive levels.
Step 4: Export Results (step4_result_assignments.py)
- Input: DataFrame + assignments from Step 3
- Output:
unified_taxonomy_full.csv- all records with unified categoriesunified_taxonomy_hierarchy.csv- unique taxonomy structure
Usage
Use scripts/pipeline.py to run the complete 4-step workflow.
See scripts/pipeline.py for:
- Complete implementation of all 4 steps
- Example code for processing multiple sources
- Command-line interface
- Individual step usage (for advanced control)
Related Skills
Agent-Reach
87.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.7k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
85.1k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
