Hnet
Association ruled based networks using graphical Hypergeometric Networks.
Install / Use
npx skills add erdogant/hnetInstalls into whichever agent you are using.
README
Key Features
| Feature | Description | Docs | Medium | Gumroad+Podcast| |---------|-------------|------|------|-------| | Association Learning | Discover significant associations across variables using statistical inference. | Link | Link | Link | | Mixed Data Handling | Works with continuous, discrete, categorical, and nested variables without heavy preprocessing. | Link | - | - | | Summarization | Summarize complex networks into interpretable structures. | Link | - | - | | Feature Importance | Rank variables by importance within associations. | Link | - | - | | Interactive Visualizations | Explore results with dynamic dashboards and d3-based visualizations. | Dashboard | - | Titanic Example | | Performance Evaluation | Compare accuracy with Bayesian association learning and benchmarks. | Link | - | - | | Interactive Dashboard | No data leaves your machine. All computations are performed locally. | Link | - | - |
Resources and Links
- Example Notebooks: Examples
- Medium Blogs Medium
- Gumroad Blogs with podcast: GumRoad
- Documentation: Website
- Bug Reports and Feature Requests: GitHub Issues
- Article: arXiv
- Article: PDF
Background
-
HNet stands for graphical Hypergeometric Networks, which is a method where associations across variables are tested for significance by statistical inference. The aim is to determine a network with significant associations that can shed light on the complex relationships across variables. Input datasets can range from generic dataframes to nested data structures with lists, missing values and enumerations.
-
Real-world data often contain measurements with both continuous and discrete values. Despite the availability of many libraries, data sets with mixed data types require intensive pre-processing steps, and it remains a challenge to describe the relationships between variables. The data understanding phase is crucial to the data-mining process, however, without making any assumptions on the data, the search space is super-exponential in the number of variables. A thorough data understanding phase is therefore not common practice.
-
Graphical hypergeometric networks (
HNet), a method to test associations across variables for significance using statistical inference. The aim is to determine a network using only the significant associations in order to shed light on the complex relationships across variables. HNet processes raw unstructured data sets and outputs a network that consists of (partially) directed or undirected edges between the nodes (i.e., variables). To evaluate the accuracy of HNet, we used well known data sets and generated data sets with known ground truth. In addition, the performance of HNet is compared to Bayesian association learning. -
HNet showed high accuracy and performance in the detection of node links. In the case of the Alarm data set we can demonstrate on average an MCC score of 0.33 + 0.0002 (P<1x10-6), whereas Bayesian association learning resulted in an average MCC score of 0.52 + 0.006 (P<1x10-11), and randomly assigning edges resulted in a MCC score of 0.004 + 0.0003 (P=0.49). HNet overcomes processes raw unstructured data sets, it allows analysis of mixed data types, it easily scales up in number of variables, and allows detailed examination of the detected associations.
Installation
Install hnet from PyPI
pip install hnet
Install from Github source
pip install git+https://github.com/erdogant/hnet
Imort Library
import hnet
print(hnet.__version__)
# Import library
from hnet import hnet
<hr>
Installation
- Install hnet from PyPI (recommended).
pip install -U hnet
Examples
- Simple example for the Titanic data set
# Initialize hnet with default settings
from hnet import hnet
# Load example dataset
df = hnet.import_example('titanic')
# Print to screen
print(df)
# PassengerId Survived Pclass ... Fare Cabin Embarked
# 0 1 0 3 ... 7.2500 NaN S
# 1 2 1 1 ... 71.2833 C85 C
# 2 3 1 3 ... 7.9250 NaN S
# 3 4 1 1 ... 53.1000 C123 S
# 4 5 0 3 ... 8.0500 NaN S
# .. ... ... ... ... ... ... ...
# 886 887 0 2 ... 13.0000 NaN S
# 887 888 1 1 ... 30.0000 B42 S
# 888 889 0 3 ... 23.4500 NaN S
# 889 890 1 1 ... 30.0000 C148 C
# 890 891 0 3 ... 7.7500 NaN Q
<a href="https://erdogant.github.io/docs/d3graph/titanic_example/index.html">Play with the interactive Titanic results.</a>
<link rel="import" href="https://erdogant.github.io/docs/d3graph/titanic_example/index.html">Example: Learn association learning on the titanic dataset
<p align="left"> <a href="https://erdogant.github.io/hnet/pages/html/Examples.html#titanic-dataset"> <img src="https://github.com/erdogant/hnet/blob/master/docs/figs/fig4.png" width="900" /> </a> </p>Example: Summarize results
Networks can become giant hairballs and heatmaps unreadable. You may want to see the general associations between the categories, instead of the label-associations. With the summarize functionality, the results will be summarized towards categories.
<p align="left"> <a href="https://erdogant.github.io/hnet/pages/html/Use%20Cases.html#summarize-results"> <img src="https://github.com/erdogant/hnet/blob/masRelated Skills
mcp
Use the `mcp_perplexity-ask_perplexity_search` tools to answer questions. You should use this instead of the `web_search` tool because it is a lot more accurate.
practical-power-systems-synthesis
This skill enables synthesis in the domain of power-systems (engineering). It represents research-level-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform synthesis operations related to power-systems.
semi-supervised-optogenetics-testing
This skill enables testing in the domain of optogenetics (neuroscience). It represents intermediate-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform testing operations related to optogenetics.
data-mining-interpretation-fundamental
This skill enables interpretation in the domain of data-mining (data-science). It represents fundamental-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform interpretation operations related to data-mining.
