Baseline Establishment

SkillDev tools

SOTA Performance Baseline Campaign — 5 strategies for systematically

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Baseline Establishment skill

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/baseline-establishment/SKILL.md and read by ahel’s review.

Strategy Routing

User IntentRoute To
Find all methods for a taskmethod-inventory
Extract scores from papersperformance-extraction
Normalize conditions across paperscondition-standardization
Check reproducibility / discrepanciesdiscrepancy-analysis
Track progress over time / headroomprogress-quantification

Manifest

Strategies (5)

StrategyPurpose
method-inventoryComprehensively identify all relevant methods for a task
performance-extractionSystematically extract performance data and conditions from papers
condition-standardizationStandardize evaluation condition differences across papers
discrepancy-analysisIdentify discrepancies between reported and reproducible scores
progress-quantificationTrack performance progress over time, quantify remaining headroom

Tactics (3)

TacticPurpose
leaderboard-harvestingSystematically collect performance data from platforms and papers
condition-normalizationCompare and standardize experimental conditions across papers
progress-curve-constructionBuild performance-over-time progress curves

Subagent SOPs (10)

SOPPurpose
method-discoveryIdentify methods via literature, leaderboards, citation chains
score-extractionExtract (Task, Dataset, Metric, Score, Conditions) tuples
condition-catalogingRecord evaluation conditions per method
reproducibility-checklist-auditAssess paper against ML Reproducibility Checklist
performance-table-assemblyAssemble unified comparison table
compute-normalizationNormalize results by compute budget
discrepancy-identificationCompare same-method scores across sources
headroom-estimationEstimate ceiling vs current SOTA gap
progress-curve-fittingConstruct performance-over-time data
baseline-synthesisProduce final structured baseline report

Budget Table

StrategyMethodsData PointsWeb Searches
method-inventory50060
performance-extraction3015040
condition-standardization206030
discrepancy-analysis154530
progress-quantification3010040
TOTAL145355200

MCP Tools

MCP ServerTools
brave-searchbrave_web_search, brave_llm_context
apifyrag-web-browser, google-scholar-scraper
alphaxivget_paper_content, answer_pdf_queries
semantic-scholarss_paper, ss_relevance_search, ss_citations, ss_references

Context Management

Campaign outputs are accumulated in the calling knowledge-acquisition context:

  • methods_inventory.json — All discovered methods with metadata
  • performance_data.json — Extracted scores with provenance
  • conditions_matrix.json — Standardized conditions per method
  • discrepancy_report.json — Flagged score inconsistencies
  • progress_curves.json — Time-series performance data
  • baseline_report.md — Final synthesized baseline document

Available Strategies

Optional, no fixed order; the final leaf is always a sop.

StrategyWhen to use
condition-standardizationStandardize evaluation condition differences across papers — 20 methods, 60 data points, 30 web searches budget
discrepancy-analysisIdentify discrepancies between reported and reproducible scores — 15 methods, 45 data points, 30 web searches budget
method-inventoryComprehensively identify all relevant methods for a task — 50 methods, 60 web searches budget
performance-extractionSystematically extract performance data and conditions from papers — 30 methods, 150 data points, 40 web searches budget
progress-quantificationTrack performance progress over time, quantify remaining headroom — 30 methods, 100 data points, 40 web searches budget

Available SOPs

Optional, no fixed order; the final leaf is always a sop.

SOPWhen to use
context-checkpointAppend research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase.
context-initCreate a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed.

Signals

GitHub stars
469
Forks
37
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
baseline-establishment
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine