degenerate-input-filtering
Filter degenerate, uninformative inputs before statistical tests: single-sequence alignments, empty files, constant features, zero-variance inputs, all-NaN columns. See nan-safe-correlation for NaN-aware correlation; statistical-analysis for test guidance.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill degenerate-input-filteringInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of degenerate-input-filtering
degenerate-input-filtering scores 91/100 on our quality scale, 1172nd of 4,619 Development & Engineering skills we index (top 26%).
Its SKILL.md is 14 KB long, well organised into 21 sections with 6 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so degenerate-input-filtering is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
degenerate-input-filtering compared with similar skills
All 4 of these similar skills score higher than degenerate-input-filtering; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| degenerate-input-filtering (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.0k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install degenerate-input-filtering?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill degenerate-input-filtering. The install tabs above show the steps for each supported agent. - Which AI agents does degenerate-input-filtering work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is degenerate-input-filtering safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is degenerate-input-filtering still maintained?
- The repository was last updated 37 days ago, so degenerate-input-filtering is actively maintained.
Skill content
View source on GitHubname: degenerate-input-filtering description: "Filter degenerate, uninformative inputs before statistical tests: single-sequence alignments, empty files, constant features, zero-variance inputs, all-NaN columns. See nan-safe-correlation for NaN-aware correlation; statistical-analysis for test guidance." license: CC-BY-4.0
Degenerate Input Filtering Guide
Overview
Degenerate inputs are data points that carry no statistical information: constant-value features, all-NaN columns, single-sequence alignments, empty files, and similar edge cases. When these reach a statistical test or model, the result is meaningless -- a correlation of NaN, a p-value of 1.0, a score of 0.0, or an outright crash. This guide establishes the mandatory practice of detecting and removing such inputs before any analysis, and of reporting every removal so that downstream consumers know the effective sample size.
Key Concepts
What Counts as Degenerate
A data point is degenerate when it cannot contribute to the statistic being computed. The root cause is always the same: the input lacks the variation or completeness that the method requires.
| Type | Example | Why It Fails | |---|---|---| | Constant-value feature | Gene with identical expression across all samples | Variance = 0; correlation, t-test, fold-change are all undefined | | All-NaN feature | Column with no valid observations | Every aggregation returns NaN | | Single-sequence alignment | BLAST result with one sequence | Score = 0.0; no pairwise comparison is possible | | Empty file | 0-byte FASTA or CSV | Parser crashes or returns an empty frame | | Single value after grouping | One sample in a treatment group | Within-group variance is undefined; group comparison is meaningless | | Zero-length sequence | Empty FASTA entry | Alignment and k-mer tools fail or produce nonsense | | Near-constant feature | Gene with one outlier and N-1 identical values | Technically non-zero variance but correlation is dominated by a single point |
Why Silent Failures Are Dangerous
Many numerical libraries do not raise errors on degenerate input. Instead they return sentinel values:
numpy.corrcoefreturnsnanfor constant columns without warning.scipy.stats.spearmanrreturns(nan, nan)when one array is constant.scipy.stats.ttest_indreturns(nan, nan)for zero-variance groups.
These NaN values propagate silently through pipelines, contaminating aggregated statistics, heatmaps, volcano plots, and ranked gene lists. By the time a researcher notices the problem, it may be unclear which upstream step introduced the NaN.
The Reporting Obligation
Filtering without reporting is nearly as harmful as not filtering. When items are silently dropped, the effective sample or feature count differs from what the user expects, leading to incorrect power calculations and misleading summary statistics. Every filtering step must print:
- The count of items removed.
- The reason for removal.
- The count of items remaining.
Decision Framework
Use this tree to determine which filtering checks apply to your data before running a statistical analysis:
Is the input tabular (DataFrame)?
├── Yes
│ ├── Are there columns with zero unique values (all NaN)? → Remove, report count
│ ├── Are there columns with exactly one unique value (constant)? → Remove, report count
│ ├── Are there columns with fewer than N valid observations? → Remove, report count
│ └── Are there near-constant columns (1 outlier, rest identical)? → Flag or remove
└── No (file list, sequence set, etc.)
├── Are there empty files (0 bytes)? → Skip, report count
├── Are there single-entry collections (e.g., 1-sequence alignment)? → Skip, report count
└── Are there zero-length entries within a file? → Filter entries, report count
| Scenario | Recommended Check | Threshold | |---|---|---| | Gene expression matrix before DE analysis | Remove zero-variance genes, remove genes detected in fewer than N samples | N = max(3, 10% of samples) | | Correlation analysis between two feature sets | Remove features constant in either set; require >= 10 shared non-NaN pairs | nunique >= 2 in both vectors; shared observations >= 10 | | Multiple sequence alignment scoring | Skip alignments with < 2 sequences | Sequence count >= 2 | | Survival analysis with grouped covariates | Remove groups with < 2 events | Events per group >= 2 | | PCA or clustering on a feature matrix | Remove zero-variance and all-NaN features | Variance > 0 and at least 1 non-NaN value | | Batch correction (e.g., ComBat) | Remove features absent in any batch | Feature detected in all batches |
Best Practices
-
Filter before any statistical computation, not after. Degenerate inputs can cause division-by-zero, NaN propagation, and inflated multiple-testing corrections. Filtering after the test means the damage is already done.
-
Use a single reusable filtering function. Centralizing the logic prevents inconsistencies between scripts and ensures the reporting format is uniform. See the reference implementation in the Workflow section below.
-
Print counts at every filtering step. Even when the count is zero, printing "0 items removed" confirms that the check ran. This makes logs self-documenting and auditable.
-
Set domain-appropriate thresholds rather than relying on generic defaults. A gene expression analysis might require detection in at least 10% of samples, while a proteomics dataset with more missingness might use 5%. Document the threshold and the rationale.
-
Treat near-constant features with caution. A feature with one non-identical value technically has non-zero variance, but its correlation with anything is driven entirely by that single point. Consider flagging these separately rather than including them blindly.
-
Never filter silently. Dropping data without reporting is a form of hidden data manipulation. Every removal must be logged with the reason and the count, so that the effective sample size is always transparent.
-
Re-check after transformations. Log-transforming expression data can introduce negative infinities from zero counts. Merging datasets can introduce new NaN columns. Run the degenerate check again after any transformation or merge step.
Common Pitfalls
-
Running a correlation on constant-valued columns and getting NaN without realizing it. The NaN result silently propagates into downstream rankings, heatmaps, or enrichment inputs, producing misleading figures.
- How to avoid: Always check
nunique() >= 2for every column before computing correlations. Use the reference implementation below.
- How to avoid: Always check
-
Scoring single-sequence alignments and reporting a score of 0.0 as a real result. Alignment scores require at least two sequences; a single-sequence "alignment" is not a failed alignment, it is a non-alignment.
- How to avoid: Count sequences before scoring. Skip and report any alignment file with fewer than 2 sequences.
-
Filtering data but not reporting the count, leading to confusion about effective sample size. A volcano plot based on 15,000 genes looks very different from one based on 18,500, but without a log message, no one knows which it is.
- How to avoid: Always print three things: count removed, reason, count remaining. Include this in every notebook and script.
-
Applying a minimum-samples threshold that is too low, retaining genes detected in only 1-2 samples. These genes produce unstable variance estimates and unreliable p-values that inflate the false discovery rate.
- How to avoid: Use
max(3, int(0.1 * n_samples))as a floor. Adjust upward for small studies where even 10% is fewer than 3.
- How to avoid: Use
-
Forgetting to re-check after a merge or transformation step. Joining two datasets can introduce all-NaN columns from non-overlapping features. Log-transforming zeros creates negative infinities. Both are degenerate.
- How to avoid: Run the degenerate filter again after every merge, join, or mathematical transformation.
-
Using
df.var() > 0without handling NaN correctly. If a column is all NaN,var()returns NaN, which is not> 0, so the column is dropped -- but the reason logged is "zero variance" rather than "all NaN," which is misleading.- How to avoid: Check for all-NaN columns separately before checking for zero variance. Report each category with its own label.
-
Assuming the input is clean because it came from a curated database. Public databases contain placeholder values, missing entries, and withdrawn records. Always validate, even when the source is trusted.
- How to avoid: Treat every input as untrusted. Run the full degenerate-input check regardless of the data source.
Workflow
The following sequential process should be applied before any statistical analysis.
-
Step 1: Load and inspect
- Load the data into a DataFrame or equivalent structure.
- Print the shape: rows, columns, and dtype summary.
-
Step 2: Remove all-NaN features
- Identify columns (or rows, depending on orientation) where every value is NaN.
- Remove them and print the count.
-
Step 3: Remove constant-value features
- Identify columns with
nunique() <= 1(after excluding NaN). - Remove them and print the count.
- Identify columns with
-
Step 4: Remove features with too few valid observations
- Set a threshold (e.g., at least 10% of samples or at least 3).
- Remove features below the threshold and print the count.
-
Step 5: Apply domain-specific checks
- For sequence data: skip single-sequence alignments, empty files, zero-length entries.
- For expression data: filter low-detection genes.
- For correlation inputs: verify both vectors are non-constant and share enough observations.
-
Step 6: Report summary
- Print a final summary: total input, total removed (by category), total remaining.
Reference Implementation
import pandas as pd
import numpy as np
def filter_degenerate(data, context="features"):
"""Filter degenerate data points and report what was removed.
Always call this BEFORE statistical analysis.
"""
n_before = len(data)
# Filter based on data type
if isinstance(data, pd.DataFrame):
# Remove constant columns (zero variance)
non_constant = data.loc[:, data.nunique() > 1]
# Remove all-NaN columns
non_nan = non_constant.dropna(axis=1, how='all')
filtered = non_nan
elif isinstance(data, list):
# Remove None, empty, and zero-length items
filtered = [x for x in data if x is not None and len(x) > 0]
else:
filtered = data
n_after = len(filtered) if not isinstance(filtered, pd.DataFrame) else filtered.shape[1]
n_removed = n_before if not isinstance(data, pd.DataFrame) else data.shape[1]
n_removed = n_removed - n_after
# MANDATORY: Print the count
print(f"Degenerate {context} filtered: {n_removed} / {n_before} removed")
print(f"Remaining {context}: {n_after}")
return filtered
Sequence Alignment Filtering
# Filter single-sequence alignments (score=0.0 is meaningless)
alignment_scores = {}
for aln_file in alignment_files:
n_seqs = count_sequences(aln_file)
if n_seqs < 2:
print(f"Skipping {aln_file}: only {n_seqs} sequence(s)")
continue
score = compute_alignment_score(aln_file)
alignment_scores[aln_file] = score
print(f"Filtered {len(alignment_files) - len(alignment_scores)} single-sequence alignments")
Gene Expression Filtering
# Filter genes with zero variance (constant expression)
gene_vars = expression_df.var(axis=1)
zero_var = (gene_vars == 0).sum()
print(f"Genes with zero variance: {zero_var}")
expression_filtered = expression_df.loc[gene_vars > 0]
# Filter genes detected in too few samples
min_samples = max(3, int(0.1 * expression_df.shape[1])) # At least 10% of samples
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
45.0kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
