Experiment Runner

SkillMonitoring & ops

Implement and execute approved research Experiment, Tuning, or Design Cards with reproducible logs, metrics, environment, and artifacts. Use for bounded coding, debugging, and runs; not for changing claims, priorities, budgets, or the Research Spine.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Experiment Runner skill

What this skill tells your AI

The instructions your AI receives, as published by cpt13-g/build-research-evidence in .agents/skills/experiment-runner/SKILL.md and read by ahel’s review.

Read research/research_state.yaml, the approved Card, research/experiment_contract.yaml, docs/PROTOCOLS.md, and docs/ROLE_HANDOFFS.md before execution. For Cheap Evaluation or refinement, also read docs/METHOD_REALIZATION_PROTOCOL.md. Refuse work before Gate 1 or a missing, draft, stale, budget-ineligible, dependency-blocked, phase-ineligible, priority-ineligible Card, and return the exact blocking fields.

Implement shared behavior in src/ and experiment differences in configs/. Preserve original data and checkpoints. Use tools/runner.py to emit the run manifest with command, Card/config identity, declared data and seed protocol, code/environment identity, timestamps, logs, raw metrics, exit status, artifact hashes, and explicit missing declarations. Do not describe a run as reproducible when its manifest is partial.

Engineering Debug and tuning are allowed within the Card. If a change affects a scientific mechanism, stop and request an approved Design Iteration. Never alter the Research Spine, Claim/EQ, priority, fairness protocol, budget or stop condition.

For cheap_evaluation, implement exactly the Card's proxy reductions and shared protocol across Candidates; label outputs preliminary and never promote them to paper evidence. For refinement_validation, modify only the approved Design Iteration dimensions and preserve the controlled comparison.

Stop once the Card completion contract is decided. Extra runs require an Orchestrator-approved named evidence gap; general completeness is insufficient. Return execution facts without a scientific verdict or the Orchestrator's desired interpretation. A failed run remains in experiment memory and goes to result-auditor when interpretable.

An Evidence Completion Queue entry is a recommendation, not execution authorization. Run it only after the Orchestrator creates and approves a normal Card and Registry activation succeeds. P3/SKIP entries are never default execution work.

Signals

GitHub stars
20
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
experiment-runner
Source
github.com/cpt13-g/build-research-evidence