Baseline Comparison Audit — is the comparison complete, fair, and significant?

SkillDev tools

Audit whether a paper's baseline comparisons are COMPLETE, FAIR, and SIGNIFICANT: a required recent SOTA baseline is missing while 'best/SOTA' is claimed (HP-MISSING-BASELINE); a baseline is undertuned / given less compute-tuning-data, run at a mismatched config, or the equal-budget ablation-as-baseline is absent (HP-WEAK-BASELINE); 'outperforms' is asserted over overlapping error bars or with no variance/seeds (HP-SIG-OVERLAP); and a cross-row 'improves over baseline by X%' is arithmetically wrong (HP-DELTA-ERROR, cross-row form only). A versioned per-domain baseline profile + a live leaderboard/recency search are assembled by the EXECUTOR as structured facts; a fresh cross-model reviewer (gpt-5.6-sol xhigh, read-only, fresh thread per dimension) PROPOSES findings, each span-anchored to a ledger claim_id; tools/adjudicate_findings.py DECIDES the verdict. Works at L0 (stated comparisons) and deepens at L2 (configs/result files). A completeness question it cannot settle internally becomes needs_external_check, never a guessed missing baseline. Emits baseline-comparison-audit.findings.json; computes NO verdict. Detect-only. Triggers: \"baseline audit\", \"missing baselines\", \"is the comparison fair\", \"weak baseline\", \"baseline 误报\", \"SOTA earned?\".

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Baseline Comparison Audit — is the comparison complete, fair, and significant? skill

What this skill tells your AI

The instructions your AI receives, as published by wanshuiyin/anti-autoresearch in skills/baseline-comparison-audit/SKILL.md and read by ahel’s review.

Audit baseline-comparison integrity for: $ARGUMENTS (requires claims.json from /evidence-ledger). Emit span-anchored baseline-comparison-audit.findings.json. This skill computes no verdict.

🔒 Do not wrap this skill in /loop, /schedule, or CronCreate. It is verdict-bearing input — it proposes the findings the deterministic adjudicator turns into the report. Re-firing it on a wall-clock timer adds no signal: its output changes only when the paper / ledger (or the live leaderboard it cross-checks) changes, not with the clock. Schedule the external wait that precedes it — ledger built → audit once. (Mirrors ARIS's external-cadence doctrine.)

Adapted from ARIS paper-claim-audit — its scope-overclaim and delta-arithmetic checks, reframed from "paper vs result files" to "is the SOTA claim earned, and is the comparison a fair fight?" — plus a per-domain baseline profile and a completeness / fairness / significance split. A favourite autoresearch shortcut is to claim SOTA while omitting the obvious recent baseline, to beat an undertuned one, or to write "outperforms" over error bars that overlap. This skill is the constraint that asks for the fair fight, pointed at a third party's submission, and it stays honest about what it cannot settle from a PDF.

Why this exists

An autoresearch pipeline (or rushed human) optimises for the headline and treats the comparison table as scaffolding to fill, not a fair experiment to run. The repeatable failure modes:

  • Completeness — "achieves state-of-the-art on GSM8K" while the obvious recent baseline a 2024–2026 reviewer expects is simply absent from the table, or the strong classical floor (BM25 for retrieval, GBDT for tabular, a linear/naive forecaster for time-series) is skipped while only weak neural baselines are beaten. HP-MISSING-BASELINE
  • Fairness — the proposed method is tuned for 100 epochs / 5 seeds / extra data, the baseline is run at default settings for 10; or the compared rows use different backbones, splits, or eval protocols; or the single most informative baseline — the method's own backbone with the new component removed, at an identical budget — is missing. HP-WEAK-BASELINE
  • Significance — "consistently outperforms" on a 0.3-point gap with overlapping error bars, with no variance / no seed count reported at all, or resting on a single dataset too thin for the "consistent / across-the-board" wording. HP-SIG-OVERLAP
  • Delta arithmetic — "improves over the strongest baseline by 16%" when the baseline row is 73.1 and the proposed row is 78.0 (+6.7% relative / +4.9 points), the two operands sitting in different cells so the single-sentence deterministic pass cannot pair them. HP-DELTA-ERROR (cross-row form)

None of these is inherently misconduct — they are what an optimizing agent does when nothing forces a fair comparison. The stated version is decidable at L0 from the manuscript; the verified version (real configs, real seeds) deepens at L2. What this skill will not do is guess: where no domain profile exists and the leaderboard search is inconclusive, the completeness question is handed off as needs_external_check, not invented.

Core principle

Ledger-anchored, span-verified, reviewer≠adjudicator, honest about what it cannot settle. Four properties:

  1. Anchor to a PAPER claim. Every above-info finding cites a ledger claim_id and quotes a verbatim span of that claim's text_span (references/integrity-forensics-contract.md rules 1–2). The "outperforms / SOTA / best / first" language lives in comparison and scope claims; the reported baseline set lives in baseline claims; values/table rows in number / table_cell claims. The anchor is whichever paper claim the finding undermines — the expected-baseline list, a leaderboard URL, or a config file:line are forensic context for the description, never the anchor.
  2. The executor assembles facts; the reviewer judges. The profile + a live WebSearch/WebFetch give a candidate expected-baseline set with sources and dates; the executor passes it as structured input and never pre-declares "baseline X is missing" (references/reviewer-independence.md). The model proposes; tools/adjudicate_findings.py decides. This skill computes no verdict.
  3. Unsettleable completeness → hand off, don't guess. "Is this really SOTA / first / the right baseline set?" cannot be closed from inside the paper. Unless an omission is unambiguous, sourced, same-benchmark, and pre-dating, emit verdict_local: needs_external_check + requires_external_check: true, not a flag (contract rule 6).
  4. Observability caps severity. Stated-comparison checks are L0; a fairness finding that needs the actual config/seed files is observability_level_required: 2 and is marked as needing L2 on a PDF-only run (references/observability-levels.md).

How this differs from the other auditors (route correctly)

AuditorQuestion it answersLevel
consistency-auditDoes the paper contradict ITSELF / described method = evaluated method? (owns text-only HP-SCOPE-INFLATE + single-sentence HP-DELTA-ERROR)L0
experiment-forensicsAre the reported numbers what the code actually computes? (fake GT, self-norm, phantom)L2
baseline-comparison-audit (this)Are the right baselines present (completeness), fairly tuned/configured (fairness), and is "outperforms/SOTA" statistically earned (significance)?L0 stated / L2 verified
citation-forensicsDo the cited baseline papers exist and support the claim?L0
presentation-signalsSurface "AI-flavor" hints (auxiliary, surface-class)L0
adversarial-case-builderStrongest evidence-bound rejection memo (no verdict weight)any

Do NOT raise here (hand off instead): generic in-text scope inflation ("comprehensive / extensive / robust" decoupled from a SOTA/comparison claim) → consistency-audit owns HP-SCOPE-INFLATE; a single-sentence "from A to B, X%" delta whose operands and stated value sit in one sentence → already caught deterministically by consistency-audit (do not re-emit — Step 5 dedups); whether a baseline number matches the repo/code → experiment-forensics (L2); whether a cited baseline paper exists / is used in-context → citation-forensics; surface / AI-flavor → presentation-signals. This skill never emits an F-pattern.

Per-domain baseline profile (PROFILE_VERSION = 0.1 — a SEED prior, always verified live)

The expected baseline set a competent 2024–2026 reviewer carries into the table. It is advisory and deliberately at the level of families / floors (not pinned method names that go stale); the live WebSearch/WebFetch leaderboard check (Step 2) is the authoritative cross-check — the profile only seeds the question. The Fairness control column names the matched-budget axis a HP-WEAK-BASELINE finding turns on; the Variance norm column is what HP-SIG-OVERLAP turns on.

Domain / benchmarkExpected baseline families (bold = the easy-to-skip floor)Fairness control (matched-budget axis)Variance norm (significance)
LLM reasoning / QA — GSM8K, MATH, MMLU, BBH, GPQAa current frontier model (Llama-3.x, Qwen2.5, DeepSeek) + the prior method on the same benchmark; strong CoT / self-consistency on the same basesame base model; identical #shots, decoding (temp / SC samples), tool access, finetune datavariance over prompts/seeds for small gaps
Image classification — ImageNet-1ka recent strong backbone at matched params/FLOPs (ConvNeXt-V2, DeiT-III, Swin-V2, MAE-ViT); a well-tuned modern CNNparams, FLOPs, input res, epochs, augmentation, pretrain datasingle run common; ±std if pretraining differs
Detection / segmentation — COCO, ADE20Ka recent strong detector/segmenter at the same backbone & schedule (DINO, Co-DETR, ViTDet, Mask2Former); a strong one-stage baselinebackbone, schedule (1×/3×), input scale, extra datasingle run common; ±std on mIoU if available
Machine translation — WMTtuned Transformer-big + a recent NMT/LLM-MT system; report COMET, not only BLEUdata, model size, beam, vocab; same test split + (de)tok protocolbootstrap CI on BLEU/COMET
Generation — FID on ImageNet/COCO, GenEvala recent strong generator at matched sampling budget (DiT, EDM2, U-ViT, LDM); report precision/recall, not only FIDNFE/sampler, params, guidance; identical FID protocol (#samples, ref stats)FID over a fixed sample size; seed/sample noise
Retrieval / RAG — BEIR, MTEB, NQa strong dense retriever + prior SOTA; BM25 (the lexical floor)same corpus, index, eval protocol (full vs sampled negatives)per-query bootstrap CI
Tabular learninga strong recent DL-tab model + prior SOTA; a well-tuned GBDT (XGBoost/LightGBM/CatBoost) — it MUST be tunedHPO-budget parity, features, splitsstd over folds/seeds
Time-series forecastinga strong recent forecaster + prior SOTA; a linear / naive-seasonal baselinelookback, horizon, normalization, splitsstd over windows/seeds
RL — control (MuJoCo/DMC), offline (D4RL), Atarituned SAC/TD3/PPO (online), CQL/IQL/Decision-Transformer (offline), Rainbow/IQN/DrQ/SPR (Atari); a well-tuned standard algorithmenv steps / frames / dataset, net size, #eval seeds & episodes≥5 seeds + std/IQM (rliable CI)
Code generation — HumanEval, MBPP, LiveCodeBencha current frontier code LLM + prior SOTA + a same-size open base; the base model w/o the proposed scaffoldmodel size, #shots, decoding, contamination windowvariance over samples (pass@k seeds)
Speech ASR — LibriSpeecha Whisper-class / Conformer system + prior SOTAtraining data, decoding / LMWER ±CI if available
Graph — OGBa strong GNN family + the prior OGB-leaderboard entryfeatures, splitsstd over seeds

Cross-domain control (always applicable, even off-profile): the single most informative baseline is the proposed method's own backbone / base model with the new component removed, run at an identical budget — the ablation-as-baseline. Its absence, or an unequal budget for it, is the most common fairness failure and is checkable for any paper, profile row or not.

Honesty rule (load-bearing): no profile row + inconclusive search ⇒ NO guessed "missing baseline". Run the fairness + significance checks (which need no profile) and emit the completeness question as needs_external_check. The profile seeds a question, never a detector.

Constants & Reviewer Calling Convention

REVIEWER_MODEL        = gpt-5.6-sol                  # different family from executor (Claude)
REVIEWER_REASONING    = xhigh                    # always; effort never lowers reviewer quality
REVIEWER_SANDBOX      = read-only                # detect-only; never mutate the paper
REVIEWER_CWD          = <paper-dir>              # so it can read claims.json + sources directly
THREAD_POLICY         = fresh mcp__codex__codex per DIMENSION (and per entry on fan-out);
                        NEVER mcp__codex__codex-reply across dimensions/entries
TAXONOMY_VERSION      = 0.5                      # references/hack-pattern-taxonomy.md
PROFILE_VERSION       = 0.1                      # the per-domain baseline profile above (advisory)
PATTERNS_OWNED        = HP-MISSING-BASELINE, HP-WEAK-BASELINE, HP-SIG-OVERLAP,
                        HP-DELTA-ERROR (cross-row comparison form only — see Step 4),
                        HP-RESOURCE-IDENTITY-MISMATCH (named dataset/model/benchmark vs its
                        public record — HF card / Papers-with-Code; gather-facts-then-judge,
                        observability_level_required 0; FP-suppress subset/variant/version)
FINDINGS_FILE         = baseline-comparison-audit.findings.json
FINDING_ID_NAMESPACE  = BC###                    # distinct from F###/NUM###/HL### (consistency), EF### (experiment)
TRACE_POLICY          = forensic (never silently dropped)
TRACE_DIR             = .aris/traces/baseline-comparison-audit/<YYYY-MM-DD>_run<NN>/
  • Executor (Claude) builds none of the judgment: it locates the ledger, extracts the comparison surface, assembles the candidate expected set with sources and publication dates, at L2 gathers mechanical config/result facts (grep/hash — listing what exists is a fact, not a judgment), passes paths + the ledger + those facts + the checklist to the reviewer, validates the reviewer's spans, and writes the findings file. It never summarizes the paper, pre-judges "X is missing", or leaks an opinion into the prompt (reviewer-independence.md). Passing what a public leaderboard says (with its date) is the same allowed division experiment-forensics (grep/hash facts) and citation-forensics (canonical metadata) use — reference facts, not hunches about the manuscript.
  • Reviewer (codex / gpt-5.6-sol) reads claims.json and the sources, decides which comparisons are incomplete / unfair / not significant, applies the known false-positive cases, and self-reports false_positive_risk. It is the evidence-extractor, not the judge.
  • Fresh thread per dimension. Completeness (Step 3) and fairness + significance + delta (Step 4) are separate fresh mcp__codex__codex calls. On — effort: max or many comparison rows, fan each comparison entry out into its own fresh call — never codex-reply carrying one entry's conclusion into another (the bias guard). codex-reply is intentionally absent from allowed-tools.

Step 0 — Preconditions: locate the ledger, read the run level

The ledger is the only structure this skill reasons over. Resolve it and read the observability level L and paper_id it was built at (each Bash block is self-contained — shell state does not persist between calls, so re-derive paths every block):

ROOT=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
# $ARGUMENTS is a paper-dir OR a claims.json path:
LEDGER="$ARGUMENTS"; [ -d "$LEDGER" ] && LEDGER="$LEDGER/claims.json"
# Only the NO-ARGUMENT case defaults to the CWD ledger. An EXPLICIT argument that
# resolves to a missing claims.json must NOT silently fall back to $(pwd) — that
# could audit the wrong paper; let the NO_LEDGER check below fire instead.
[ -z "$ARGUMENTS" ] && LEDGER="$(pwd)/claims.json"
python3 - "$LEDGER" <<'PY'
import json, sys, os, collections
p = sys.argv[1]
if not os.path.isfile(p):
    sys.exit("NO_LEDGER: claims.json not found. Run /evidence-ledger FIRST "
             "(it writes artifact_manifest.json + claims.json).")
d = json.load(open(p, encoding="utf-8"))
claims = d.get("claims", [])
by = collections.Counter(c.get("type") for c in claims)
print("LEDGER       =", os.path.abspath(p))
print("PAPER_DIR    =", os.path.dirname(os.path.abspath(p)) or ".")
print("PAPER_ID     =", d.get("paper_id", "?"))
print("RUN_LEVEL_L  =", d.get("observability_level", 0))
print("CLAIMS       =", len(claims), dict(by))
# applicability signal — comparison / scope / baseline / table_cell are this skill's inputs:
rel = sum(by.get(t, 0) for t in ("comparison", "scope", "baseline", "table_cell"))
print("APPLICABLE   =", "yes" if rel else "low (no comparison/scope/baseline/table claims)")
PY

Failure handling. If NO_LEDGER is printed, stop and tell the user to run /evidence-ledger first — this skill never re-reads the raw PDF and invents its own structure (contract rule 1). Carry L, PAPER_ID, and the absolute LEDGER / PAPER_DIR into every step below.

Step 1 — Extract the comparison surface from the ledger (decide whether to run)

Pull the claims this audit reasons over and decide if there is anything to audit:

ROOT=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
LEDGER="<abs path to claims.json from Step 0>"
python3 - "$LEDGER" <<'PY'
import json, re, sys, collections
d = json.load(open(sys.argv[1], encoding="utf-8"))
claims = d.get("claims", [])
COMPARE = re.compile(r"\b(state[- ]of[- ]the[- ]art|SOTA|outperform\w*|best|"
                     r"surpass\w*|beats?|superior|first to|consistently|"
                     r"compared?\s+(?:to|with)|baseline|prior\s+(?:work|art))\b", re.I)
anchors, baselines = [], []
for c in claims:
    t, span = c.get("type"), c.get("text_span", "")
    # the SOTA / outperforms assertion (the anchor for completeness + fairness):
    if t == "comparison" or (t == "scope" and COMPARE.search(span)):
        anchors.append((c["claim_id"], t, c.get("location", {}).get("section", "?"), span[:150]))
    # the baseline SET the paper reports (mechanical; the reviewer decides completeness):
    if t == "baseline":
        baselines.append((c["claim_id"], c.get("location", {}).get("section", "?"), span[:150]))
print(f"ANCHOR (comparison/SOTA) claims: {len(anchors)}   baseline-list claims: {len(baselines)}")
for cid, t, sec, sp in anchors[:40]:
    print(f"  [anchor:{t}] {cid} [{sec}] {sp!r}")
for cid, sec, sp in baselines[:20]:
    print(f"  [baseline] {cid} [{sec}] {sp!r}")
nums = collections.Counter(c.get("type") for c in claims if c.get("type") in ("number", "table_cell"))
mets = collections.Counter((c.get("value") or {}).get("metric") for c in claims
                           if (c.get("value") or {}).get("metric"))
print("VALUE CLAIMS =", dict(nums), "  METRICS (seed the profile row) =", dict(mets))
print("APPLICABLE   =", "yes" if (anchors or baselines) else "no -> write [] and stop")
PY

Branch. If APPLICABLE = no (zero comparison/SOTA/baseline claims), this skill is not applicable: write an empty baseline-comparison-audit.findings.json ([]), record a one-line NOT_APPLICABLE reason in the trace (Step 7), and stop. Silent skip is forbidden — the orchestrator globs *.findings.json and expects the file to exist. Otherwise record the anchor claims (the SOTA/comparison assertions) and the reported-baseline list for the prompts. A purely mechanical grep helps surface the table/baseline names (do not judge completeness here — that is the reviewer's job):

LEDGER="<abs path to claims.json from Step 0>"
grep -rInE '\\begin\{tabular|\\caption|baseline|w\.r\.t|vs\.?|\bours?\b' \
    "$(dirname "$LEDGER")" --include='*.tex' 2>/dev/null | head -60

Step 2 — Assemble the candidate expected-baseline set (profile + live search + recency guard)

Determine the benchmark/task from the anchor claims (and the method section), then build a candidate expected set with sources — structured evidence, not a verdict. The recency guard is what stops you from naming a hallucinated or concurrent baseline as "missing":

  1. Profile lookup. If the task is in the per-domain profile above, take its expected baseline families (incl. the floor) and the matched-budget axis. If no row matches → mark the domain NO_PROFILE; completeness defaults to needs_external_check.
  2. Live leaderboard cross-check (the authoritative source; record the query + the top systems + their dates/venues):
    WebSearch: "<benchmark> state-of-the-art <paper/current year> leaderboard"
    WebSearch: "<benchmark> papers with code"
    WebFetch:  <the Papers-with-Code / leaderboard URL>   # top 5–8 systems + their dates
    
  3. Recency + existence guard. For each candidate baseline, record its same_benchmark? · published_before_paper? · source_url · date. A system concurrent with or post-dating the audited paper is a legitimate omission (a false positive for "missing"), not a flag.

Create the run's trace dir now — its first use is the file written just below, so it must exist before Step 7. Reuse this exact RUNDIR in Steps 3–7 (do not create a second one):

DATE=$(date +%Y-%m-%d); N=1
while [ -d ".aris/traces/baseline-comparison-audit/${DATE}_run$(printf %02d $N)" ]; do N=$((N+1)); done
RUNDIR=".aris/traces/baseline-comparison-audit/${DATE}_run$(printf %02d $N)"; mkdir -p "$RUNDIR"
echo "RUNDIR = $RUNDIR"   # carry this exact path forward (shell state does not persist)

Write these facts (not opinions) into $RUNDIR/expected_baseline_set.json:

{
  "task": "GSM8K grade-school math (LLM reasoning)",
  "profile_version": "0.1",
  "profile_hit": true,
  "leaderboard_source": "https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k  (read <UTC date>)",
  "expected_baselines": [
    {"name": "Self-Consistency CoT", "year": 2023, "same_benchmark": true, "published_before_paper": true, "source": "<url>"},
    {"name": "<recent frontier model> few-shot CoT", "year": 2024, "same_benchmark": true, "published_before_paper": true, "source": "<url>"}
  ],
  "notes": "off-profile or inconclusive search -> completeness becomes needs_external_check, not a guess"
}

Failure handling. No network / search fails → fall back to the profile alone and mark every completeness candidate requires_external_check: true (you could not confirm recency/availability). An empty candidate set + NO_PROFILE ⇒ skip the completeness flag and emit a single needs_external_check info finding instead. Never phrase the set as "the paper is missing X" — that is the reviewer's call.

Step 3 — Completeness pass (cross-model, fresh thread) → HP-MISSING-BASELINE

Open a fresh mcp__codex__codex thread (Reviewer Calling Convention). The reviewer reads claims.json from its cwd for the present baselines and compares against your external expected-set facts; every finding anchors to a ledger claim_id. Send EXACTLY (fill every [ ... ]):

mcp__codex__codex:
  model: gpt-5.6-sol
  config: {"model_reasoning_effort": "xhigh"}
  sandbox: read-only
  cwd: <absolute PAPER_DIR from Step 0>
  prompt: |
    You are a baseline-COMPLETENESS forensics reviewer. You judge ONE thing: given
    what this paper claims ("state-of-the-art / best / first / outperforms prior
    work") on a benchmark, is an OBVIOUS, RECENT, RELEVANT baseline absent from the
    comparison? You do NOT judge whether numbers are real and you do NOT grade the
    paper. Describe a discrepancy to CHECK, never an accusation; hand off what you
    cannot ground.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
152
Forks
8
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
baseline-comparison-audit
Source
github.com/wanshuiyin/anti-autoresearch