compare-evaluation-protocols

SkillDev tools

Lets your agent compare evaluation protocols and judge which differences could change benchmark results.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the compare-evaluation-protocols skill

About this skill

Build a protocol-difference matrix and estimate which differences can materially change measured performance.

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/compare-evaluation-protocols/SKILL.md and read by ahel’s review.

Purpose

Build a protocol-difference matrix and estimate which differences can materially change measured performance.

Input contract

required: [protocol_records, metric_schema, comparison_target]
optional: [paired_results, sensitivity_assumptions]
constraints: [differences require explicit protocol fields and a comparable outcome]

Procedure

  1. Extract protocol elements into a normalized comparison schema.
  2. Align datasets, populations, metrics, baselines, and evaluation conditions.
  3. Mark differences and assess their plausible performance effect.
  4. Separate observed effects from unresolved protocol confounding.

Output contract

produces: [protocol_difference_matrix, materiality_assessment, confounding_notes, comparability_judgment]
delta_fields: [findings, evidence_updates, uncertainties, open_questions]

Quality gates

  • The compared metric and target are held constant or explicitly qualified.
  • Missing protocol fields remain visible.

Failure and counterexamples

Do not attribute score differences to method quality when protocol differences are unmeasured.

Provenance map

  • resolved: knowledge-acquisition-evaluation-protocol-comparison

Signals

GitHub stars
501
Forks
41
Last commit
Sep 2026
Advanced
Catalog kind
skill
Key
compare-evaluation-protocols
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine