Strategy: Comparison Design

SkillMedia

Design fair comparison experiments against baselines and competing methods

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Strategy: Comparison Design skill

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/comparison-design/SKILL.md and read by ahel’s review.

Question: How much better is our method than the baseline?

Methodology

  • Fair Comparison Protocol (Bouthillier 2021): Control all confounds, same compute budget, same tuning effort.
  • Multi-Baseline Comparison: Compare against multiple baselines (SOTA, simple, ablated).
  • Multi-Dataset Evaluation: Test across diverse datasets to avoid dataset-specific overfitting.
  • Bayesian Comparison (Benavoli 2017): Posterior probability of superiority, not just p-values.
  • Bootstrap/Permutation Tests: Non-parametric significance without distributional assumptions.

Execution Flow

  1. baseline-selection → Select appropriate baselines (SOTA, simple, oracle)
  2. metric-specification → Define primary metric and secondary metrics
  3. sample-size-estimation → Power analysis for detecting meaningful differences
  4. seed-protocol-design → Ensure fair random initialization across methods
  5. environment-specification → Lock environment to prevent confounds
  6. reproducibility-protocol (tactic) → Ensure all results are reproducible
  7. statistical-method-selection (tactic) → Choose Bayesian or frequentist comparison

Budget Gate

Comparison ScopeBaselinesDatasetsSeedsMin Runs
Minimal1 SOTA + 1 simple136
Standard2-3 baselines2-3530-45
Comprehensive4+ baselines3-55-10100+
Publication-readyAll relevant5+10+200+

Available Tactics

Optional, no fixed order; the final leaf is always a sop.

TacticWhen to use
reproducibility-protocolEnsure experiment reproducibility through systematic environment and seed control
statistical-method-selectionSelect appropriate statistical methods for experiment analysis

Available SOPs

Optional, no fixed order; the final leaf is always a sop.

SOPWhen to use
baseline-selectionSelect appropriate baselines for experimental comparison
environment-specificationSOP: define complete experiment environment specification
metric-specificationDefine experiment metrics and significance standards
sample-size-estimationSOP: power analysis and required experiment count estimation
seed-protocol-designSOP: design random seed strategy for reproducibility

Signals

GitHub stars
469
Forks
37
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
comparison-design
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine