hypogenic-hypothesis-generation
LLM-driven hypothesis generation/testing on tabular data. Three methods: HypoGeniC (data-driven), HypoRefine (literature+data), Union. Iterative refinement, Redis caching, multi-hypothesis inference. Manual: hypothesis-generation; ideation: scientific-brainstorming.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill hypogenic-hypothesis-generationInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of hypogenic-hypothesis-generation
hypogenic-hypothesis-generation scores 91/100 on our quality scale, 1102nd of 2,866 Automation skills we index (top 39%).
Its SKILL.md is 15 KB long, well organised into 62 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so hypogenic-hypothesis-generation is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
hypogenic-hypothesis-generation compared with similar skills
All 4 of these similar skills score higher than hypogenic-hypothesis-generation; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| hypogenic-hypothesis-generation (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| rufloby ruvnet | 100 | 73.9k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install hypogenic-hypothesis-generation?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill hypogenic-hypothesis-generation. The install tabs above show the steps for each supported agent. - Which AI agents does hypogenic-hypothesis-generation work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is hypogenic-hypothesis-generation safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is hypogenic-hypothesis-generation still maintained?
- The repository was last updated 37 days ago, so hypogenic-hypothesis-generation is actively maintained.
Skill content
View source on GitHubname: "hypogenic-hypothesis-generation" description: "LLM-driven hypothesis generation/testing on tabular data. Three methods: HypoGeniC (data-driven), HypoRefine (literature+data), Union. Iterative refinement, Redis caching, multi-hypothesis inference. Manual: hypothesis-generation; ideation: scientific-brainstorming." license: "MIT"
HypoGeniC Hypothesis Generation
Overview
HypoGeniC automates scientific hypothesis generation and testing using LLMs on tabular datasets. Given labeled data (e.g., deception detection, AI-content identification), it generates testable hypotheses, iteratively refines them against validation performance, and runs inference to classify new samples. It supports three approaches: purely data-driven (HypoGeniC), literature-integrated (HypoRefine), and mechanistic union of both.
When to Use
- Generating testable hypotheses from labeled observational datasets without prior theory
- Systematically testing multiple competing hypotheses on empirical data
- Combining insights from research papers with data-driven pattern discovery
- Accelerating hypothesis ideation in domains like deception detection, content analysis, mental health indicators
- Benchmarking LLM-based hypothesis generation methods against few-shot baselines
- For manual hypothesis formulation frameworks, use hypothesis-generation knowhow
- For general-purpose ML classification without hypothesis interpretability, use scikit-learn-machine-learning
Prerequisites
- Python packages:
hypogenic - Optional: Redis server (port 6832) for LLM response caching; GROBID for PDF literature processing
- API keys: OpenAI, Anthropic, or compatible LLM API key in environment
- Data: Labeled JSON datasets in HypoGeniC format (see Key Concepts)
pip install hypogenic
# Optional: clone example datasets
git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data
git clone https://github.com/ChicagoHAI/Hypothesis-agent-datasets.git ./data_lit
Quick Start
from hypogenic import BaseTask
import re
# Custom label extractor (must match dataset label format)
def extract_label(text: str) -> str:
match = re.search(r'final answer:\s+(.*)', text, re.IGNORECASE)
return match.group(1).strip() if match else text.strip()
# 1. Load task from config
task = BaseTask(
config_path="./data/your_task/config.yaml",
extract_label=extract_label
)
# 2. Generate hypotheses (data-driven)
task.generate_hypotheses(
method="hypogenic",
num_hypotheses=20,
output_path="./output/hypotheses.json"
)
# 3. Run inference on test set
results = task.inference(
hypothesis_bank="./output/hypotheses.json",
test_data="./data/your_task/your_task_test.json"
)
print(f"Accuracy: {results['accuracy']:.3f}")
Workflow
Step 1: Prepare Dataset
Create train/val/test JSON files with text features and labels.
import json
# Dataset: each key maps to a list of equal length
dataset = {
"headline_1": [
"What Up, Comet? You Just Got *PROBED*",
"Scientists Made a Breakthrough in Quantum Computing"
],
"headline_2": [
"Scientists Were Holding Their Breath Today. Here's Why.",
"New Quantum Computer Achieves Milestone"
],
"label": [
"Headline 2 has more clicks than Headline 1",
"Headline 1 has more clicks than Headline 2"
]
}
# All lists must have equal length; labels must match extract_label output
for split in ["train", "val", "test"]:
with open(f"my_task_{split}.json", "w") as f:
json.dump(dataset, f, indent=2)
print(f"Created dataset with {len(dataset['label'])} samples")
Step 2: Create Task Configuration
Write a config.yaml defining dataset paths and prompt templates.
# config.yaml structure (write as YAML file)
config = """
task_name: my_task
train_data_path: ./my_task_train.json
val_data_path: ./my_task_val.json
test_data_path: ./my_task_test.json
prompt_templates:
observations: |
Feature 1: ${text_features_1}
Feature 2: ${text_features_2}
Observation: ${label}
batched_generation:
system: "You are a research scientist generating hypotheses."
user: "Generate ${num_hypotheses} testable hypotheses from these observations."
inference:
system: "You are evaluating a hypothesis against data."
user: "Hypothesis: ${hypothesis}\\nSample: ${sample_text}\\nFinal answer: ${label}"
is_relevant:
system: "Check hypothesis relevance."
user: "Is this hypothesis relevant? ${hypothesis}"
"""
with open("config.yaml", "w") as f:
f.write(config)
print("Configuration written to config.yaml")
Step 3: Implement Label Extraction
Define a custom extract_label function matching your label format.
import re
def extract_label(llm_output: str) -> str:
"""Parse LLM output to extract predicted label.
Must return labels matching the 'label' field values in the dataset.
Default: searches for 'final answer: <label>' pattern.
"""
match = re.search(r'final answer:\s+(.*)', llm_output, re.IGNORECASE)
if match:
return match.group(1).strip()
# Domain-specific fallback
if "Final prediction:" in llm_output:
return llm_output.split("Final prediction:")[-1].strip()
return llm_output.strip()
# Test against expected labels
assert extract_label("Final answer: Headline 1") == "Headline 1"
print("Label extractor validated")
Step 4: Generate Hypotheses (HypoGeniC)
Run data-driven hypothesis generation with iterative refinement.
from hypogenic import BaseTask
task = BaseTask(
config_path="./config.yaml",
extract_label=extract_label
)
# Generate hypotheses: initializes from data subset, iteratively refines
task.generate_hypotheses(
method="hypogenic", # Data-driven generation
num_hypotheses=20, # Target number of hypotheses
output_path="./output/hypotheses.json"
)
# CLI equivalent:
# hypogenic_generation --config config.yaml --method hypogenic --num_hypotheses 20
print("Hypothesis bank saved to ./output/hypotheses.json")
Step 5: Run Inference
Test generated hypotheses against the test set.
results = task.inference(
hypothesis_bank="./output/hypotheses.json",
test_data="./my_task_test.json"
)
print(f"Test accuracy: {results['accuracy']:.3f}")
print(f"Predictions: {results['predictions'][:5]}")
# CLI equivalent:
# hypogenic_inference --config config.yaml --hypotheses output/hypotheses.json
Step 6: Literature-Integrated Generation (HypoRefine)
Combine literature insights with data-driven hypotheses.
# Requires GROBID setup and preprocessed PDFs
# bash ./modules/setup_grobid.sh # first time
# bash ./modules/run_grobid.sh # start GROBID service
# python pdf_preprocess.py --task_name my_task
task.generate_hypotheses(
method="hyporefine",
num_hypotheses=15,
literature_path="./literature/my_task/",
output_path="./output/"
)
# Generates 3 hypothesis banks:
# - HypoRefine (integrated literature+data)
# - Literature-only hypotheses
# - Literature union HypoRefine
print("HypoRefine generation complete: 3 hypothesis banks created")
Step 7: Multi-Hypothesis Inference
Test multiple hypotheses simultaneously for ensemble classification.
from examples.multi_hyp_inference import run_multi_hypothesis_inference
results = run_multi_hypothesis_inference(
config_path="./config.yaml",
hypothesis_bank="./output/hypotheses.json",
test_data="./my_task_test.json"
)
print(f"Multi-hypothesis accuracy: {results['accuracy']:.3f}")
Key Parameters
| Parameter | Default | Range / Options | Effect |
|-----------|---------|-----------------|--------|
| method | "hypogenic" | "hypogenic", "hyporefine", "union" | Generation strategy |
| num_hypotheses | 20 | 5-50 | Number of hypotheses to generate |
| batch_size | 5 | 3-10 | Samples per generation batch |
| max_iterations | 10 | 1-50 | Refinement iterations |
| temperature | 0.7 | 0.0-1.0 | LLM sampling temperature |
| confidence_threshold | 0.7 | 0.5-0.95 | Inference confidence cutoff |
| num_papers | 10 | 5-30 | Papers for HypoRefine literature extraction |
| inference_method | "voting" | "voting", "weighted", "ensemble" | How multiple hypotheses combine predictions |
Key Concepts
Dataset Format
HypoGeniC expects JSON files with parallel lists:
{
"text_features_1": ["sample_1_feat1", "sample_2_feat1"],
"text_features_2": ["sample_1_feat2", "sample_2_feat2"],
"label": ["class_A", "class_B"]
}
- All lists must have equal length
- Feature keys are customizable (
review_text,post_content, etc.) - Labels must match the
extract_label()output format exactly - Three splits required:
<TASK>_train.json,<TASK>_val.json,<TASK>_test.json
Three Generation Methods
| Method | Input | Process | Best For | |--------|-------|---------|----------| | HypoGeniC | Data only | Init from subset, iteratively refine on validation | Exploratory research, novel datasets without literature | | HypoRefine | Data + PDFs | Extract literature insights, merge with data patterns, refine both | Extending or validating existing theories | | Union | Literature + HypoGeniC | Mechanistic combination, deduplication | Maximum hypothesis diversity and coverage |
Configuration Template
Minimal required config.yaml structure:
task_name: my_task
train_data_path: ./my_task_train.json
val_data_path: ./my_task_val.json
test_data_path: ./my_task_test.json
model:
name: "gpt-4" # or claude-3, gpt-3.5-turbo
api_key_env: "OPENAI_API_KEY"
temperature: 0.7
generation:
method: "hypogenic"
num_hypotheses: 20
batch_size: 5
max_iterations: 10
cache:
enabled: true # Redis on localhost:6832
host: "localhost"
port: 6832
prompt_templates:
observations: |
Feature 1: ${text_features_1}
Observation: ${label}
batched_generation:
system: "Generate testable hypotheses."
user: "Generate ${num_hypotheses} hypotheses."
inference:
system: "Evaluate hypothesis against sample."
user: "Hypothesis: ${hypothesis}\nSample: ${sample_text}"
is_relevant:
system: "Check relevance."
user: "Is ${hypothesis} relevant?"
Common Recipes
Recipe: Custom Task from Scratch
When to use: creating a new classification task with domain-specific data.
import json
from hypogenic import BaseTask
# 1. Prepare data splits
for split_name, data in [("train", train_data), ("val", val_data), ("test", test_data)]:
with open(f"my_task_{split_name}.json", "w") as f:
json.dump(data, f)
# 2. Define domain-specific label extractor
def my_extractor(text):
if "positive" in text.lower():
return "positive"
elif "negative" in text.lower():
return "negative"
return text.strip()
# 3. Create task and run full pipeline
task = BaseTask(config_path="./my_task/config.yaml", extract_label=my_extractor)
task.generate_hypotheses(method="hypogenic", num_hypotheses=15, output_path="./output/")
results = task.inference(hypothesis_bank="./output/hypotheses.json")
print(f"Custom task accuracy: {results['accuracy']:.3f}")
Recipe: Literature Processing Setup
When to use: setting up GROBID for PDF-to-structured-text conversion before HypoRefine.
# 1. Setup GROBID (first time only)
bash ./modules/setup_grobid.sh
# 2. Place PDFs in literature directory
mkdir -p literature/my_task/raw/
cp papers/*.pdf literature/my_task/raw/
# 3. Start GROBID and process
bash ./modules/run_grobid.sh
cd examples && python pdf_preprocess.py --task_name my_task
# Output: structured text files in literature/my_task/processed/
Recipe: Union Method for Maximum Coverage
When to use: combining literature and data-driven hypotheses for comprehensive coverage.
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
ruflo
73.9k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
