Result Auditor
SkillProductivityValidate and scientifically interpret completed research runs, including negative, null, boundary, confounded, failed, and inconclusive outcomes. Use after execution to assess evidence and claim impact; not to edit methods or schedule experiments.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Result Auditor skill
What this skill tells your AI
The instructions your AI receives, as published by cpt13-g/build-research-evidence in .agents/skills/result-auditor/SKILL.md and read by ahel’s review.
Read research/research_state.yaml, the Card, contract, Runner manifest, raw metrics/logs, docs/PROTOCOLS.md, and docs/ROLE_HANDOFFS.md. For Candidate or refinement results, also read docs/METHOD_REALIZATION_PROTOCOL.md; for post-P0 aggregation read docs/EVIDENCE_COMPLETION_PROTOCOL.md. Perform the first pass without the Orchestrator's interpretation or preferred story. Fix the initial validity, fairness and result verdict before reconciliation. Record the reviewed inputs, deterministic checks and any isolation exception. Never interpret an execution failure as evidence against the scientific hypothesis.
Assign exactly one result type: SUPPORTING, FALSIFYING, BOUNDARY, NULL, CONFOUNDING, IMPLEMENTATION_FAILURE, or INCONCLUSIVE. Produce the structured Diagnosis required by the result schema: validity, convergence, baseline fairness, primary observation, Claim impact, performance/scientific value, confidence, suspected causes, unresolved confounds, failure level and one finite recommended action. Check seed/statistical protocol, data leakage and missing artifacts.
Use F1 only for execution failures and recommend Debug/Tuning. Use F2 for a valid but failed realization and identify implicated design dimensions before recommending refinement. Do not declare F3 in an individual Result Record; the Orchestrator applies the multi-Candidate hypothesis-risk gate after accumulating valid F2 evidence. A NULL result must retain explicit scientific value. Cheap results remain PRELIMINARY_NON_PAPER regardless of Metric.
Write a schema-valid result record linked to the immutable run manifest and preserve meaningful outcomes in experiment memory. Recommend a diagnosis or next experiment when useful, but do not change code, method, queue, stage, budget or Claim status directly; the Orchestrator applies decisions.
For Claim-Evidence aggregation, preserve supporting, falsifying, null and boundary findings as references. Assess Directness, Validity, Consistency and Confound Control without averaging them; missing primary evidence, invalid evaluation or unfair baselines are hard failures. Cheap and Tuning results remain outside formal evidence.
Signals
- GitHub stars
- 20
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
result-auditor- Source
- github.com/cpt13-g/build-research-evidence