Benchmark Archaeology

SkillDev tools

Evaluation Methodology Archaeology Campaign — 5 strategies for systematic

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Benchmark Archaeology skill

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/benchmark-archaeology/SKILL.md and read by ahel’s review.

Systematic excavation and critical analysis of AI/ML evaluation methodology. Treats benchmarks as historical artifacts requiring forensic examination — uncovering hidden assumptions, methodological drift, validity decay, and coverage gaps that accumulate over time.

Strategy Routing

SignalRoute To
"audit this benchmark", "benchmark quality", "BetterBench"benchmark-audit
"saturation", "ceiling", "score plateau", "when will X be solved"saturation-analysis
"does it actually measure", "construct validity", "what does score mean"validity-probing
"what's not tested", "coverage gaps", "missing capabilities"coverage-mapping
"different papers get different scores", "protocol differences"protocol-forensics

Manifest

Strategies (5)

StrategyPurpose
benchmark-auditSystematic quality assessment using BetterBench 46-criterion framework
saturation-analysisTrack score trajectories, detect saturation and failure points
validity-probingChallenge construct validity — does benchmark measure claimed capability?
coverage-mappingMap evaluation coverage, identify untested capability dimensions
protocol-forensicsAnalyze evaluation protocol differences across papers for same benchmark

Tactics (3)

TacticPurpose
score-trajectory-analysisCollect historical scores, fit saturation curves, detect inflection points
artifact-detectionDetect annotation artifacts and shortcuts in benchmarks
evaluation-protocol-comparisonCompare implementation differences of same benchmark across papers

Subagent SOPs (9 + 1 shared)

SOPPurpose
benchmark-inventoryIdentify and catalog all relevant benchmarks in target domain
metric-decompositionDecompose composite metrics into constituent signals
contamination-auditDetect train-test data leakage and memorization artifacts
construct-validity-assessmentEvaluate whether benchmark measures its claimed capability
documentation-auditAssess documentation completeness against BetterBench/Datasheets standards
capability-taxonomy-mappingBuild capability taxonomy, map existing benchmark coverage
leaderboard-dynamics-analysisAnalyze leaderboard score distributions, compression, selective reporting
protocol-element-extractionExtract evaluation protocol parameters from papers
benchmark-synthesisProduce final structured audit report
saturation-detection (shared)Detect saturation signals in score trajectories (from literature-survey)

Budget Table

StrategyBenchmarksPapersWeb Searches
benchmark-audit53040
saturation-analysis155060
validity-probing34030
coverage-mapping203050
protocol-forensics56030
Total48210210

MCP Tools

MCP ServerTools
brave-searchbrave_web_search, brave_llm_context
apifyrag-web-browser, google-scholar-scraper
alphaxivget_paper_content, answer_pdf_queries
semantic-scholarss_paper, ss_relevance_search, ss_citations, ss_references

Context Management

All outputs write to context/benchmark-archaeology/:

context/benchmark-archaeology/
  audit/              # benchmark-audit outputs
  saturation/         # saturation-analysis outputs
  validity/           # validity-probing outputs
  coverage/           # coverage-mapping outputs
  forensics/          # protocol-forensics outputs
  synthesis/          # Final cross-strategy synthesis

Each strategy maintains its own state ledger within its output directory.

Available Strategies

Optional, no fixed order; the final leaf is always a sop.

StrategyWhen to use
benchmark-auditSystematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches
coverage-mappingMap evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches
protocol-forensicsAnalyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches
saturation-analysisTrack score trajectories, detect saturation/failure points — 15 benchmarks, 50 papers, 60 web searches
validity-probingChallenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches

Available SOPs

Optional, no fixed order; the final leaf is always a sop.

SOPWhen to use
context-checkpointAppend research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase.
context-initCreate a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed.

Signals

GitHub stars
469
Forks
37
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
benchmark-archaeology
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine