dataset-quality-audit
Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions.
Install / Use
npx skills add zebbern/claude-code-guide --skill dataset-quality-auditInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Tags
Our assessment of dataset-quality-audit
dataset-quality-audit scores 91/100 on our quality scale, 75th of 340 Data & Analytics skills we index (top 23%).
Its SKILL.md is 3.9 KB long, well organised into 13 sections with 4 code examples: a solid amount of guidance for an agent.
With 4,638 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 2 days ago, so dataset-quality-audit is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
dataset-quality-audit compared with similar skills
All 4 of these similar skills score higher than dataset-quality-audit; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| dataset-quality-audit (this skill)by zebbern | 91 | 4.6k | 2d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
Frequently asked questions
- How do I install dataset-quality-audit?
- Run
npx skills add zebbern/claude-code-guide --skill dataset-quality-audit. The install tabs above show the steps for each supported agent. - Which AI agents does dataset-quality-audit work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is dataset-quality-audit safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is dataset-quality-audit still maintained?
- The repository was last updated 2 days ago, so dataset-quality-audit is actively maintained.
Skill content
View source on GitHubname: dataset-quality-audit description: "Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data." license: MIT
dataset-quality-audit
A data quality auditing tool that runs 12-dimension quality checks on tabular data, producing per-dimension scores (0–100), an overall grade, and actionable fix suggestions.
Capabilities
| Dimension | Description | |-----------|-------------| | Missing Values | Count and percentage of null/NaN values per column | | Duplicate Rows | Number and percentage of fully duplicated rows | | Type Consistency | Mixed types within a single column (e.g., numbers mixed with text) | | Value Range / Outliers | Outlier detection using the IQR method | | Format Compliance | Consistency of date, email, phone number, and other formatted fields | | Uniqueness Constraints | Whether ID-type columns contain duplicates | | Whitespace Issues | Leading/trailing spaces, empty strings, whitespace-only values | | Constant Columns | Columns with only a single unique value (zero information) | | Distribution Skewness | Whether numeric columns have excessive skewness | | Column Naming | Spaces, special characters, or inconsistent casing in column names | | Cardinality Anomalies | Unusually high or low number of unique values | | Cross-Column Consistency | Logical checks across columns (e.g., start date before end date) |
Quick Start
# Basic quality check
python3 scripts/data_quality_checker.py data.csv
# Save report as JSON
python3 scripts/data_quality_checker.py data.csv --output report.json
# Specify ID columns (for uniqueness checks)
python3 scripts/data_quality_checker.py users.csv --id-columns "user_id,email"
# Specify date columns (for format checks)
python3 scripts/data_quality_checker.py orders.csv --date-columns "created_at,updated_at"
Detailed Usage
Basic Invocation
python3 scripts/data_quality_checker.py <data-file> [options]
Parameters
| Parameter | Short | Required | Default | Description |
|-----------|-------|----------|---------|-------------|
| input | — | Yes | — | Path to input file (CSV/TSV/Excel/JSON) |
| --output | -o | No | stdout | Path for the JSON report output |
| --id-columns | -id | No | Auto-detect | Comma-separated column names that should be unique |
| --date-columns | -dc | No | Auto-detect | Comma-separated column names containing dates |
| --sample | -s | No | All rows | Number of rows to sample (useful for large files) |
| --encoding | -e | No | utf-8 | File encoding |
Output Format (JSON)
{
"file": "data.csv",
"rows": 10000,
"columns": 15,
"overall_score": 78.5,
"grade": "B",
"dimensions": {
"missing_values": {
"score": 85.0,
"issues": [
{"column": "age", "missing_count": 150, "missing_pct": 1.5, "suggestion": "Fill with median or mode"}
]
},
"duplicates": {
"score": 95.0,
"issues": [...]
}
},
"top_suggestions": [
"Column 'age' has 1.5% missing values — consider filling with the median",
"Found 200 fully duplicated rows — consider deduplication"
]
}
Grading Scale
| Grade | Score Range | Meaning | |-------|------------|---------| | A+ | 95–100 | Excellent quality — ready for use as-is | | A | 90–95 | Good quality — minor issues only | | B | 80–90 | Moderate quality — recommended to fix before use | | C | 60–80 | Poor quality — significant cleaning required | | D | 40–60 | Very poor quality — many issues need attention | | F | 0–40 | Essentially unusable — requires re-collection or major cleanup |
Dependencies
- Python 3.8+
- pandas
- numpy
pip install pandas numpy
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
