Audit a dataset
SkillDatabases & dataAudit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Audit a dataset skill
What this skill tells your AI
The instructions your AI receives, as published by pku-yuangroup/openai4s in skills/audit-dataset/SKILL.md and read by ahel’s review.
Use this skill before statistics, model training, or external publication. The goal is a compact, machine-readable audit plus explicit decisions about every issue that could invalidate downstream results.
Workflow
- Load records without silently coercing values. Preserve source row IDs.
- Call
audit_rowson a representative or complete list of row mappings. - Inspect missingness and observed type mixtures column by column.
- Resolve duplicate records and non-unique identifiers deliberately.
- If a split column exists, check both stable IDs and grouping entities for train/validation/test overlap.
- Record accepted exceptions, then rerun the audit and save the JSON result next to the cleaned dataset.
Import and run
Hyphenated Skill directories are loaded with importlib:
from importlib import import_module
audit_rows = import_module("audit-dataset.kernel").audit_rows
report = audit_rows(
rows,
target="label",
id_columns=("sample_id",),
group_columns=("patient_id",),
split_column="split",
)
rows must be a sequence of mappings. The report contains row and column
counts, per-column missing/type/unique summaries, duplicate row and ID counts,
target frequencies, and split-leakage examples.
Interpretation
- Mixed numeric/string types usually indicate parsing or sentinel-value bugs.
- Missingness is a property of both the data and the collection process; do not impute before checking whether it correlates with label, site, or time.
- Duplicate IDs are not automatically duplicate observations. Decide whether repeated measures are expected and group them during splitting.
- Any patient, molecule scaffold, time series, or near-duplicate entity shared across evaluation boundaries can inflate performance even when row IDs differ.
- A clean structural audit does not establish representativeness, label validity, causal identifiability, or ethical suitability.
Required output
Report the checks performed, blocking findings, accepted exceptions, and the exact source artifact/version. Never describe a dataset as clean without naming the leakage keys and missing-value policy that were checked.
Signals
- GitHub stars
- 404
- Forks
- 48
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
audit-dataset- Source
- github.com/pku-yuangroup/openai4s