homer-motif-analysis
De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER. findMotifsGenome.pl finds over-represented patterns vs background; annotatePeaks.pl assigns context (TSS distance, gene, repeat).
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysisInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of homer-motif-analysis
homer-motif-analysis scores 91/100 on our quality scale, 1164th of 4,619 Development & Engineering skills we index (top 26%).
Its SKILL.md is 25 KB long, well organised into 126 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so homer-motif-analysis is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
homer-motif-analysis compared with similar skills
All 4 of these similar skills score higher than homer-motif-analysis; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| homer-motif-analysis (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.0k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install homer-motif-analysis?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis. The install tabs above show the steps for each supported agent. - Which AI agents does homer-motif-analysis work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is homer-motif-analysis safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is homer-motif-analysis still maintained?
- The repository was last updated 37 days ago, so homer-motif-analysis is actively maintained.
Skill content
View source on GitHubname: "homer-motif-analysis" description: "De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER. findMotifsGenome.pl finds over-represented patterns vs background; annotatePeaks.pl assigns context (TSS distance, gene, repeat). Use after MACS3 to identify enriched TFs, annotate peaks with nearest genes, and validate ChIP-seq via the target motif." license: "GPL-3.0"
HOMER — Motif Analysis and Peak Annotation
Overview
HOMER (Hypergeometric Optimization of Motif EnRichment) is a suite of Perl/C++ tools for analyzing genomic regulatory elements. Its two primary commands are findMotifsGenome.pl, which performs de novo motif discovery and known motif enrichment against JASPAR/HOMER databases, and annotatePeaks.pl, which maps each peak to the nearest gene, distance to TSS, and genomic feature class (promoter, intron, intergenic, repeat). HOMER takes BED-format peak files from MACS3 or similar peak callers and a reference genome assembly as input, and outputs HTML/text reports ranking enriched motifs by p-value and fold enrichment over a matched background.
When to Use
- Identifying which transcription factors are bound in a ChIP-seq peak set by enriching their known motifs from JASPAR or the HOMER motif library
- Discovering novel sequence motifs de novo in open chromatin regions from ATAC-seq without prior knowledge of the binding TF
- Comparing motif landscapes between two conditions (e.g., treated vs. untreated peak sets) by running HOMER with one set as target and the other as background
- Annotating genomic peaks with nearest genes and distance to TSS for downstream functional analysis or integration with DESeq2 results
- Validating ChIP-seq experiment quality: a successful pull-down should show the target TF's canonical motif as the top hit
- Use
macs3-peak-callingfirst to generate the peak BED files that serve as input to HOMER - Use
jaspar-databaseto cross-reference HOMER-discovered motifs with JASPAR IDs and additional TF metadata - Use
MEME-CHIP(web or local) when you need a more probabilistic ZOOPS/TCM model or the MEME Suite ecosystem - Use
AME(part of MEME Suite) as a faster alternative for known motif scanning without de novo discovery
Prerequisites
- Software: HOMER (Perl + compiled binaries), conda or manual install
- Genomes: must download genome sequence via
installGenome.plafter HOMER install - Input: BED file of peaks (at minimum: chr, start, end columns); ideally summit-centered peaks from MACS3
- Python packages (for parsing/visualization):
pandas,matplotlib,seaborn
Check before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v findMotifsGenome.plfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run findMotifsGenome.plrather than barefindMotifsGenome.pl.
# Install HOMER via conda (recommended — handles Perl dependencies)
conda install -c bioconda homer
# Verify installation
findMotifsGenome.pl 2>&1 | head -3
# Usage: findMotifsGenome.pl <peak/BED file> <genome> <output directory> [options]
annotatePeaks.pl 2>&1 | head -3
# Usage: annotatePeaks.pl <peak/BED file> <genome> [options]
# Install reference genomes (downloads 2-way masker + sequence; ~3–10 GB each)
installGenome.pl hg38
installGenome.pl mm10
# Install Python parsing dependencies
pip install pandas matplotlib seaborn
Pre-flight Interview
Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: genomeAssembly
kind: derived
source: upstream
ask: "Which assembly are the peak coordinates on?"
default: "carried from the peak-calling stage"
- id: D2
param: searchWindow
kind: required
source: user
ask: "How wide a window around each peak centre should be searched?"
default: "200 bp for TF ChIP, 150 bp for ATAC"
- id: D3
param: backgroundRegions
kind: required
source: user
ask: "Compare peaks against GC-matched random genomic regions, or against a specific control set?"
default: "auto-generated GC-matched background"
- id: D4
param: repeatMasking
kind: required
source: user
ask: "Mask repetitive sequence, which otherwise dominates the enrichment as false positives?"
default: "masked"
- id: D5
param: denovoDiscovery
kind: required
source: user
ask: "Search for previously undescribed motifs, or only test the known-motif library?"
default: "known motifs only - de novo discovery is roughly ten times slower"
- id: D6
param: motifLengths
kind: optional
source: user
depends_on: [D5]
ask: "Which motif widths should the de novo search try?"
default: "8, 10, 12"
skip_if: "de novo discovery disabled"
- id: D7
param: numDenovoMotifs
kind: optional
source: user
depends_on: [D5]
ask: "How many de novo motifs should be reported?"
default: 25
skip_if: "de novo discovery disabled"
- id: D8
param: knownMotifMismatches
kind: optional
source: user
ask: "How loosely may a sequence match a known motif and still count?"
default: 2
- id: D9
param: threads
kind: never_ask
source: data
reason: "Affects runtime only, not the enrichment"
default: "min(8, available_cores)"
D1 is derived rather than asked: peak coordinates are meaningless outside the assembly they were called on, so the answer is whatever the upstream stage used, not a preference. D3 is the decision most often left at its default without thought - a background that does not match the peaks in GC content produces enrichment for GC content rather than for biology.
Quick Start
# Run de novo + known motif enrichment on TF ChIP-seq peaks (hg38, 200 bp window)
findMotifsGenome.pl peaks/tf_chip_summits.bed hg38 motif_output/ \
-size 200 -mask -p 4
# Annotate peaks with nearest genes and genomic features
annotatePeaks.pl peaks/tf_chip_peaks.narrowPeak hg38 > annotated_peaks.txt
echo "Top known motif:"
head -2 motif_output/knownResults.txt | tail -1 | cut -f1-4
echo "Annotated peaks: $(wc -l < annotated_peaks.txt) lines"
Workflow
Step 1: Installation and Genome Setup
Install HOMER and download the reference genome sequence required for motif analysis.
# Activate conda environment (or use existing env)
conda create -n homer_env -c bioconda homer python=3.10 -y
conda activate homer_env
# List available genomes
installGenome.pl list
# Install human (hg38) and mouse (mm10) genomes
# Downloads masked genome sequence and annotation files
installGenome.pl hg38
# Output: Installing hg38... Done. (3-5 min, ~3 GB)
installGenome.pl mm10
# Output: Installing mm10... Done. (3-5 min, ~2.5 GB)
# Verify genome is installed
ls ~/.homer/data/genomes/hg38/
# genome.fa chrom.sizes ...
# Check HOMER motif database
ls ~/.homer/data/knownTFs/
# vertebrates.motifs jaspar.motifs ...
Step 2: Prepare Input Peak File
Prepare a summit-centered BED file from MACS3 output for optimal motif resolution.
# Option A: Use MACS3 summit file directly (already 1 bp summit positions)
# Expand summits to ±100 bp (200 bp total) centered on summit
awk 'BEGIN{OFS="\t"} {
start = ($2 - 100 < 0) ? 0 : $2 - 100;
print $1, start, $2 + 100, $4, $5
}' peaks/tf_chip_summits.bed > peaks/tf_chip_200bp.bed
echo "Summit-centered peaks: $(wc -l < peaks/tf_chip_200bp.bed)"
# Summit-centered peaks: 12453
# Option B: Use narrowPeak file directly (HOMER accepts multi-column BED)
# HOMER uses columns 1-3 (chr, start, end) and centers internally with -size
cp peaks/tf_chip_peaks.narrowPeak peaks/input_peaks.bed
# Option C: Prepare a custom background region file (matched GC content)
# HOMER auto-generates background if not provided, but explicit background
# is recommended when comparing two peak sets
# Use the control peak set or random genomic regions as background:
bedtools shuffle -i peaks/tf_chip_peaks.narrowPeak \
-g ~/.homer/data/genomes/hg38/chrom.sizes \
-excl peaks/tf_chip_peaks.narrowPeak > peaks/background_regions.bed
echo "Background regions: $(wc -l < peaks/background_regions.bed)"
# Background regions: 12453
Step 3: De Novo Motif Discovery
Run findMotifsGenome.pl for de novo motif discovery and known motif enrichment simultaneously.
mkdir -p motif_output/
# Full run: de novo + known motif enrichment
# -size 200: use 200 bp window centered on peak midpoint
# -mask: mask repetitive elements (recommended for clean motifs)
# -p 4: use 4 CPU threads
# -S 25: find top 25 de novo motifs (default)
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_output/ \
-size 200 \
-mask \
-p 4 \
-S 25
# Check progress output:
# Reading genome sizes for hg38 ...
# Scanning for motifs...
# Optimizing 25 motifs...
# Done! Output in motif_output/
echo "Known results: $(wc -l < motif_output/knownResults.txt) motifs"
echo "De novo motifs: $(ls motif_output/homerResults/*.motif 2>/dev/null | wc -l) motifs"
# Known results: 392 motifs
# De novo motifs: 25 motifs
# For mouse peaks (mm10)
# findMotifsGenome.pl peaks/atac_peaks.bed mm10 motif_output_mm10/ \
# -size 200 -mask -p 4
Step 4: Known Motif Enrichment Only
Scan peaks for occurrences of a specific known motif or skip de novo discovery for speed.
# Skip de novo discovery (faster when you only need known motifs)
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_output_known/ \
-size 200 \
-mask \
-p 4 \
-nomotif
echo "Known motif results: $(wc -l < motif_output_known/knownResults.txt)"
# Known motif results: 392
# Find occurrences of a specific motif across peaks (outputs peak-level annotation)
# Extract the motif matrix file for the TF of interest from homerResults/
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_scan_out/ \
-size 200 \
-mask \
-find motif_output/homerResults/motif1.motif \
> peaks_with_motif1.txt
echo "Peaks containing motif1: $(wc -l < peaks_with_motif1.txt)"
# Peaks containing motif1: 8941
# Custom background: compare treated vs. control peak sets
findMotifsGenome.pl peaks/treated_peaks.bed hg38 motif_treated_vs_ctrl/ \
-size 200 \
-mask \
-p 4 \
-bg peaks/control_peaks.bed
Step 5: Peak Annotation
Use annotatePeaks.pl to assign each peak to a genomic feature and nearest gene.
# Annotate peaks with nearest gene and TSS distance
# Outputs a tab-delimited file with genomic context for each peak
annotatePeaks.pl peaks/tf_chip_peaks.narrowPeak hg38 \
> annotated_peaks.txt
echo "Annotated peaks: $(($(wc -l < annotated_peaks.txt) - 1)) peaks"
# Annotated peaks: 12453 peaks
# Preview column headers and first peak
head -2 annotated_peaks.txt | cut -f1-10
# Annotate ATAC-seq peaks (same command, different input)
annotatePeaks.pl peaks/atac_sample_peaks.narrowPeak hg38 \
> annotated_atac.txt
# Generate TSS-distance histogram (for tag density plots)
# annotatePeaks.pl can compute read density around peaks with -d flag
# annotatePeaks.pl tss hg38 -size 4000 -hist 10 \
# -d chip_tagdir/ > tss_histogram.txt
Step 6: Parse HOMER Results with Python
Read knownResults.txt and de novo motif files into pandas for downstream analysis.
import pandas as pd
import subprocess
import io
# --- Parse known motif enrichment results ---
# knownResults.txt columns:
# Motif Name | Consensus | P-value | Log P-value | q-value | # Target Seqs w/ motif | % Target | # Bg Seqs w/ motif | % Bg
known_cols = [
"motif_name", "consensus", "pvalue", "log_pvalue",
"qvalue", "n_target_seqs", "pct_target",
"n_bg_seqs", "pct_bg"
]
known = pd.read_csv(
"motif_output/knownResults.txt",
sep="\t", header=0, names=known_cols
)
# Convert string percentages to floats
known["pct_target"] = known
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
45.0kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
