establish-empirical-baseline

SkillDev tools

Lets your agent build a fair performance baseline by collecting, comparing, and normalizing results across methods.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the establish-empirical-baseline skill

About this skill

Establish a fair empirical baseline by inventorying methods, extracting comparable performance, normalizing conditions/compute, checking discrepancies, and estimating progress/headroom.

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/establish-empirical-baseline/SKILL.md and read by ahel’s review.

Purpose

Establish a fair empirical baseline by inventorying methods, extracting comparable performance, normalizing conditions/compute, checking discrepancies, and estimating progress/headroom.

Input contract

required: [method_records, benchmark_or_task, performance_measure]
optional: [historical_series, compute_metadata, condition_schema]
constraints: [comparability fields and source provenance required]

Execution protocol

Do not perform called SOP operations inline; each loaded SOP owns its contract and thresholds.

  1. You MUST load skill inventory-reference-items to inventory methods and comparison records.
  2. You MUST load skill extract-evidence-record to extract performance evidence. You MUST load skill audit-reporting-quality to audit missing or ambiguous reporting. You MUST load skill normalize-comparison-scale to normalize units, data, compute, and evaluation protocol.
  3. You MUST load skill detect-performance-discrepancy to detect comparison discrepancies. You MUST load skill estimate-performance-headroom to estimate headroom. You MUST load skill analyze-temporal-trajectory to analyze progress and leaderboard dynamics. You MUST load skill check-dominance to identify dominated and incomparable records.
  4. Synthesize the baseline with uncertainty and known incomparable records. If the resulting methods or baselines require explicit comparative selection, consider rank-candidates as the next tactic.

Output contract

produces: [method_inventory, normalized_baseline, discrepancy_report, progress_curve, headroom_estimate]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions]

Thresholds and quality gates

  • Baseline acquisition gates are relative to a declared eligible universe: record numerator, denominator, batch increment, stopping reason, and source references for each ratio.
  • method-inventory: method coverage ratio reaches a justified floor over the eligible method universe.
  • performance-extraction: comparable-record coverage ratio reaches a justified floor over eligible records.
  • condition-standardization: complete condition-vector ratio reaches a justified floor over comparable records.
  • discrepancy-analysis: score-pair coverage ratio reaches a justified floor over eligible comparison pairs.
  • progress-quantification: historical-time coverage and independent-source ratio reach justified floors; stop when added periods no longer change the trajectory conclusion.
  • Across modes, report marginal information gain and saturation state when added records or periods no longer change the baseline conclusion.
  • Normalization must expose condition, compute, metric, and unit transformations.

Failure and counterexamples

Do not call a baseline fair when conditions are missing, metrics are incomparable, or leaderboard values are copied without protocol verification. Mark headroom unknown when the historical series is below its floor.

Provenance map

9 architecture old entries: baseline, inventory, extraction, standardization, discrepancy, progress, leaderboard, normalization, curve construction. Repeated reporting prose compressed.

Legacy context checkpoint / Delta notes

Append method IDs, normalized records, excluded records with reasons, discrepancy pairs, progress model, and headroom uncertainty.

Preserved source criteria ledger

sourcesource linekindsource criterion
baseline-establishment28textual\
baseline-establishment40textual\
baseline-establishment58textual\
baseline-establishment70numeric-table\
baseline-establishment71numeric-table\
baseline-establishment72numeric-table\
baseline-establishment73numeric-table\
baseline-establishment74numeric-table\
baseline-establishment75numeric-table\
method-inventory23numeric-table\
method-inventory24numeric-table\
method-inventory25numeric-table\
method-inventory30textual
method-inventory33numeric-table\
method-inventory34numeric-table\
method-inventory35numeric-table\
method-inventory36numeric-table\
method-inventory37numeric-table\
method-inventory38textual
method-inventory41numericCannot exit until methods_discovered >= 40 (80% of target).
performance-extraction18textualExtract structured performance data from papers, leaderboards, and reproducibility studies. Each data point is a (Task, Dataset, Metric, Score, Conditions) tuple with full provenance. Prioritizes primary sources (original papers) but cross-references against leaderboards and third-party reproductions.
performance-extraction24numeric-table\
performance-extraction25numeric-table\
performance-extraction26numeric-table\
performance-extraction27numeric-table\
performance-extraction32textual
performance-extraction35numeric-table\
performance-extraction36numeric-table\
performance-extraction37numeric-table\
performance-extraction38numeric-table\
performance-extraction39numeric-table\
performance-extraction40numeric-table\
performance-extraction41textual
performance-extraction44numericCannot exit until data_points >= 120 (80% of target).
condition-standardization25numeric-table\
condition-standardization26numeric-table\
condition-standardization27numeric-table\
condition-standardization28numeric-table\
condition-standardization33textual
condition-standardization36numeric-table\
condition-standardization37numeric-table\
condition-standardization38numeric-table\
condition-standardization39numeric-table\
condition-standardization40numeric-table\
condition-standardization41textual
condition-standardization44numericCannot exit until data_points_standardized >= 48 (80% of target).
condition-standardization60textual3. Group methods by comparable condition sets
condition-standardization63textual6. Produce fair comparison subsets where conditions are controlled
discrepancy-analysis24numeric-table\
discrepancy-analysis25numeric-table\
discrepancy-analysis26numeric-table\
discrepancy-analysis27numeric-table\
discrepancy-analysis32textual
discrepancy-analysis35numeric-table\
discrepancy-analysis36numeric-table\
discrepancy-analysis37numeric-table\
discrepancy-analysis38numeric-table\
discrepancy-analysis39numeric-table\
discrepancy-analysis40textual
discrepancy-analysis43numericCannot exit until score_pairs_compared >= 36 (80% of target).
discrepancy-analysis52textual- reproducibility-checklist-audit - Assess paper reproducibility completeness
discrepancy-analysis59textual4. Apply reproducibility-checklist-audit to papers with large discrepancies
discrepancy-analysis85textual"reproducibility_checklist_score": 0,
progress-quantification26numeric-table\
progress-quantification27numeric-table\
progress-quantification28numeric-table\
progress-quantification29numeric-table\
progress-quantification34textual
progress-quantification37numeric-table\
progress-quantification38numeric-table\
progress-quantification39numeric-table\
progress-quantification40numeric-table\
progress-quantification41numeric-table\
progress-quantification42numeric-table\
progress-quantification43textual
progress-quantification46numericCannot exit until historical_data_points >= 80 (80% of target).
leaderboard-harvesting43numeric- Flag discrepancies > 1 standard deviation
leaderboard-harvesting58textual## Minimum Yield
leaderboard-harvesting62numeric-table\
leaderboard-harvesting63numeric-table\
leaderboard-harvesting64numeric-table\
leaderboard-harvesting65numeric-table\
condition-normalization28textual- Random seeds: number of runs, seed selection, variance reported
condition-normalization51textual### Stage 4: Fair Comparison Baseline
condition-normalization53textualApply normalization to produce fair comparison subsets:
condition-normalization58textualYield: Fair comparison tables with methodology notes.
condition-normalization60textual## Minimum Yield
condition-normalization64numeric-table\
condition-normalization65numeric-table\
condition-normalization66numeric-table\
condition-normalization67numeric-table\
progress-curve-construction62textual## Minimum Yield
progress-curve-construction66numeric-table\
progress-curve-construction67numeric-table\
progress-curve-construction68numeric-table\
progress-curve-construction69numeric-table\
progress-curve-construction73numeric- progress-curve-fitting (for Stages 2-3)

Context checkpoint / Delta notes

Append method IDs, normalized records, excluded records with reasons, discrepancy pairs, progress model, and headroom uncertainty.

Signals

GitHub stars
501
Forks
41
Last commit
Sep 2026
Advanced
Catalog kind
skill
Key
establish-empirical-baseline
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine