audit-benchmark-validity
SkillMonitoring & opsLets your agent check whether a benchmark actually measures what it claims, spotting contamination and flaws.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the audit-benchmark-validity skill
About this skill
Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.
What this skill tells your AI
The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/audit-benchmark-validity/SKILL.md and read by ahel’s review.
Purpose
Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.
Input contract
required: [benchmark_spec, task_definition, metric_definition, evaluation_records]
optional: [leaderboard_history, protocol_versions, coverage_target]
constraints: [claims must be linked to benchmark evidence]
Execution protocol
Do not perform called SOP operations inline; each loaded SOP owns its contract and thresholds.
- You MUST load skill
inventory-reference-itemsto inventory benchmark components and versions. You MUST load skilldecompose-evaluation-metricto decompose the evaluation metric. You MUST load skillassess-construct-validityto assess the benchmark construct. You MUST load skillextract-evaluation-protocolto reconstruct protocol versions. - Select validity, contamination, saturation, coverage, protocol-forensics, or evaluation-comparison mode.
- You MUST load skill
audit-data-contaminationto audit contamination. You MUST load skillmap-coverage-spaceto map data and task coverage. You MUST load skillanalyze-leaderboard-dynamicsto analyze saturation and temporal dynamics. You MUST load skillcompare-evaluation-protocolsto compare protocol and metric variants. You MUST load skillprobe-benchmark-artifactto probe artifacts and shortcut paths. - You MUST load skill
audit-reporting-qualityto audit reporting quality and return the validity verdict with threats, evidence, and required repairs. If the audit exposes systematic coverage gaps, considercoverage-white-space-search. If the benchmark or metric may share assumptions with the tested system, consideraudit-validator-independence.
Output contract
produces: [validity_verdict, threat_register, contamination_findings, coverage_map, protocol_drift_report, repair_actions]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions, recommended_jumps]
Thresholds and quality gates
- Every source HARD-GATE remains mandatory; no exit with an untested construct, contamination path, metric pathology, coverage claim, or protocol change.
- Resource and sampling gates use relative benchmark/target/evidence coverage rather than fixed benchmark, paper, or web counts. Declare the eligible universe, record numerator, denominator, batch increment, stopping reason, and source references.
- Saturation claims require an explicit stopping criterion and evidence that additional search/testing no longer changes the conclusion; report marginal information gain and saturation state.
- BetterBench-style criterion lists remain content checklists. Their item count is not converted into a percentage; the audit reports criterion coverage ratio over the declared applicable set and an independent-source ratio.
- Evaluation comparisons must state the controlled protocol difference and its expected impact.
Failure and counterexamples
Reject "valid" when benchmark artifact probes are absent, contamination is unknown but ignored, or leaderboard gains cannot be separated from protocol drift. A high score is not evidence of construct validity by itself.
Provenance map
7 architecture old entries: archaeology, audit, saturation, validity probing, coverage mapping, protocol forensics, evaluation comparison. Provider-specific retrieval compressed.
Legacy context checkpoint / Delta notes
Append benchmark version, construct claims, probes, contamination evidence, coverage gaps, protocol diffs, verdict, and repair decisions.
Preserved source criteria ledger
| source | source line | kind | source criterion |
|---|---|---|---|
| benchmark-archaeology | 39 | numeric-table | \ |
| benchmark-archaeology | 72 | numeric-table | \ |
| benchmark-archaeology | 73 | numeric-table | \ |
| benchmark-archaeology | 74 | numeric-table | \ |
| benchmark-archaeology | 75 | numeric-table | \ |
| benchmark-archaeology | 76 | numeric-table | \ |
| benchmark-archaeology | 77 | numeric-table | \ |
| benchmark-audit | 28 | numeric-table | \ |
| benchmark-audit | 29 | numeric-table | \ |
| benchmark-audit | 30 | numeric-table | \ |
| benchmark-audit | 35 | textual | |
| benchmark-audit | 38 | numeric-table | \ |
| benchmark-audit | 39 | numeric-table | \ |
| benchmark-audit | 40 | numeric-table | \ |
| benchmark-audit | 41 | numeric-table | \ |
| benchmark-audit | 42 | numeric-table | \ |
| benchmark-audit | 43 | numeric-table | \ |
| benchmark-audit | 44 | numeric-table | \ |
| benchmark-audit | 45 | numeric-table | \ |
| benchmark-audit | 46 | textual | |
| benchmark-audit | 49 | numeric | Cannot exit until 80% of all targets met. |
| benchmark-audit | 81 | numeric | betterbench_score: float # 0-1, proportion of 46 criteria met |
| saturation-analysis | 28 | numeric-table | \ |
| saturation-analysis | 29 | numeric-table | \ |
| saturation-analysis | 30 | numeric-table | \ |
| saturation-analysis | 35 | textual | |
| saturation-analysis | 38 | numeric-table | \ |
| saturation-analysis | 39 | numeric-table | \ |
| saturation-analysis | 40 | numeric-table | \ |
| saturation-analysis | 41 | numeric-table | \ |
| saturation-analysis | 42 | numeric-table | \ |
| saturation-analysis | 43 | numeric-table | \ |
| saturation-analysis | 44 | numeric-table | \ |
| saturation-analysis | 45 | numeric-table | \ |
| saturation-analysis | 46 | textual | |
| saturation-analysis | 49 | numeric | Cannot exit until 80% of all targets met. |
| saturation-analysis | 87 | numeric | estimated_time_to_ceiling: string # e.g., "6-12 months" |
| saturation-analysis | 89 | numeric | score_compression: float # top-10 score range |
| validity-probing | 28 | numeric-table | \ |
| validity-probing | 29 | numeric-table | \ |
| validity-probing | 30 | numeric-table | \ |
| validity-probing | 35 | textual | |
| validity-probing | 38 | numeric-table | \ |
| validity-probing | 39 | numeric-table | \ |
| validity-probing | 40 | numeric-table | \ |
| validity-probing | 41 | numeric-table | \ |
| validity-probing | 42 | numeric-table | \ |
| validity-probing | 43 | numeric-table | \ |
| validity-probing | 44 | numeric-table | \ |
| validity-probing | 45 | numeric-table | \ |
| validity-probing | 46 | textual | |
| validity-probing | 49 | numeric | Cannot exit until 80% of all targets met. |
| coverage-mapping | 21 | textual | Build a comprehensive map of "what we can and cannot measure" for a given AI capability domain. Identify white spaces where important capabilities lack rigorous evaluation, and redundancies where multiple benchmarks test the same narrow skill. |
| coverage-mapping | 27 | numeric-table | \ |
| coverage-mapping | 28 | numeric-table | \ |
| coverage-mapping | 29 | numeric-table | \ |
| coverage-mapping | 34 | textual | |
| coverage-mapping | 37 | numeric-table | \ |
| coverage-mapping | 38 | numeric-table | \ |
| coverage-mapping | 39 | numeric-table | \ |
| coverage-mapping | 40 | numeric-table | \ |
| coverage-mapping | 41 | numeric-table | \ |
| coverage-mapping | 42 | numeric-table | \ |
| coverage-mapping | 43 | numeric-table | \ |
| coverage-mapping | 44 | numeric-table | \ |
| coverage-mapping | 45 | textual | |
| coverage-mapping | 48 | numeric | Cannot exit until 80% of all targets met. |
| protocol-forensics | 20 | textual | Expose the "reproducibility gap" in benchmark evaluation by documenting how papers differ in their implementation of supposedly standardized evaluation protocols. Quantify the score variance attributable to protocol differences rather than model improvements. |
| protocol-forensics | 26 | numeric-table | \ |
| protocol-forensics | 27 | numeric-table | \ |
| protocol-forensics | 28 | numeric-table | \ |
| protocol-forensics | 33 | textual | |
| protocol-forensics | 36 | numeric-table | \ |
| protocol-forensics | 37 | numeric-table | \ |
| protocol-forensics | 38 | numeric-table | \ |
| protocol-forensics | 39 | numeric-table | \ |
| protocol-forensics | 40 | numeric-table | \ |
| protocol-forensics | 41 | numeric-table | \ |
| protocol-forensics | 42 | numeric-table | \ |
| protocol-forensics | 43 | numeric-table | \ |
| protocol-forensics | 44 | textual | |
| protocol-forensics | 47 | numeric | Cannot exit until 80% of all targets met. |
| protocol-forensics | 61 | textual | 1. Target Selection: Choose 5 benchmarks with known reproducibility issues or high paper volume |
| protocol-forensics | 63 | numeric | a. Collect 10-15 papers that report results on the same benchmark |
| protocol-forensics | 76 | textual | 6. Synthesis: Produce per-benchmark forensics report with reproducibility recommendations |
| protocol-forensics | 98 | textual | reproducibility_grade: A\ |
| evaluation-protocol-comparison | 12 | textual | Compare how different papers implement the same benchmark to expose hidden protocol variance that undermines cross-paper score comparability. |
| evaluation-protocol-comparison | 18 | numeric | Collect 10-15 papers that report results on the target benchmark: |
| evaluation-protocol-comparison | 36 | numeric-table | \ |
| evaluation-protocol-comparison | 48 | textual | - Low: Minor variations (e.g., different random seeds) |
| evaluation-protocol-comparison | 82 | textual | cross_paper_comparability: high\ |
| evaluation-protocol-comparison | 91 | textual | \ |
| evaluation-protocol-comparison | 93 | numeric-table | \ |
| evaluation-protocol-comparison | 94 | numeric-table | \ |
| evaluation-protocol-comparison | 95 | numeric-table | \ |
| evaluation-protocol-comparison | 96 | numeric-table | \ |
Context checkpoint / Delta notes
Append benchmark version, construct claims, probes, contamination evidence, coverage gaps, protocol diffs, verdict, and repair decisions.
Signals
- GitHub stars
- 501
- Forks
- 41
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Key
audit-benchmark-validity- Source
- github.com/yogsoth-ai/de-anthropocentric-research-engine