Benchmark Audit Strategy

SkillDev tools

Systematic quality assessment using BetterBench 46-criterion framework

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Benchmark Audit Strategy skill

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/benchmark-audit/SKILL.md and read by ahel’s review.

Systematic quality assessment of AI/ML benchmarks using the BetterBench 46-criterion framework, Datasheets for Datasets standards, and established psychometric evaluation principles.

Purpose

Produce a structured quality report for each target benchmark covering: documentation completeness, construct validity indicators, statistical robustness, maintenance status, and known failure modes.

Budget

ResourceFloorTarget
Benchmarks audited35
Papers read2030
Web searches2540

State Ledger

<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks audited | 0 | 5 | PENDING |
| Papers fetched | 0 | 30 | PENDING |
| Papers read | 0 | 20 | PENDING |
| Web searches | 0 | 40 | PENDING |
| Documentation audits complete | 0 | 5 | PENDING |
| Metric decompositions complete | 0 | 5 | PENDING |
| Contamination checks complete | 0 | 5 | PENDING |
| Synthesis reports produced | 0 | 5 | PENDING |
</HARD-GATE>

Cannot exit until 80% of all targets met.

Available Tactics

  • artifact-detection — Probe for annotation artifacts and dataset shortcuts

Available SOPs

  • benchmark-inventory — Identify target benchmarks in domain
  • metric-decomposition — Decompose composite metrics into constituent signals
  • contamination-audit — Detect train-test data leakage
  • documentation-audit — Assess documentation completeness (BetterBench/Datasheets)
  • benchmark-synthesis — Produce final structured audit report

Execution Guidance

  1. Inventory Phase: Use benchmark-inventory to identify 5 benchmarks in target domain
  2. Per-Benchmark Loop (repeat for each benchmark): a. Gather benchmark paper, documentation, leaderboard via web searches b. Run documentation-audit against BetterBench 46 criteria c. Run metric-decomposition on primary metric(s) d. Run contamination-audit checking known training corpora e. Run artifact-detection tactic if annotation-based benchmark f. Collect findings into per-benchmark report
  3. Synthesis Phase: Run benchmark-synthesis to produce cross-benchmark comparison

Output Format

benchmark_audit:
  benchmark_name: string
  version: string
  betterbench_score: float  # 0-1, proportion of 46 criteria met
  documentation_grade: A|B|C|D|F
  metric_analysis:
    primary_metric: string
    ceiling_effects: boolean
    polarity_issues: list
  contamination_risk: low|medium|high|critical
  artifact_risk: low|medium|high
  maintenance_status: active|stale|abandoned
  key_findings: list[string]
  recommendations: list[string]

Available Tactics

Optional, no fixed order; the final leaf is always a sop.

TacticWhen to use
artifact-detectionDetect annotation artifacts and shortcuts in benchmarks

Available SOPs

Optional, no fixed order; the final leaf is always a sop.

SOPWhen to use
benchmark-synthesisProduce final structured audit report
contamination-auditDetect train-test data leakage and memorization artifacts
documentation-auditAssess documentation completeness against BetterBench/Datasheets standards
knowledge-acquisition-benchmark-inventoryIdentify and catalog all relevant benchmarks in target domain
metric-decompositionDecompose composite metrics into constituent signals, analyze polarity and ceiling effects

Signals

GitHub stars
469
Forks
37
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
benchmark-audit
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine