Experiment
SkillDev toolsUse when a quest is ready for a concrete implementation pass or a main experiment run tied to a selected idea and an accepted baseline.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Experiment skill
What this skill tells your AI
The instructions your AI receives, as published by openlair/dr-claw in skills/ds-experiment/SKILL.md and read by ahel’s review.
Use this skill for the main evidence-producing runs of the quest.
Interaction discipline
- Follow the shared interaction contract injected by the system prompt.
- For ordinary active work, prefer a concise progress update once work has crossed roughly 6 tool calls with a human-meaningful delta, and do not drift beyond roughly 12 tool calls or about 8 minutes without a user-visible update.
- Keep ordinary subtask completions concise. When a main experiment actually finishes or reaches a stage-significant checkpoint, upgrade to a richer
artifact.interact(kind='milestone', reply_mode='threaded', ...)report rather than another short progress line. - That richer experiment-stage milestone report should normally cover: what run finished, the headline result versus baseline or expectation, the main caveat, and the exact recommended next action.
- That richer milestone report is still normally non-blocking. If the next route is already justified locally, continue automatically after reporting rather than idling for acknowledgment.
- If the active communication surface is QQ and QQ milestone media is enabled in config, a completed main experiment may attach one summary PNG to that richer milestone update.
- That PNG should be a connector-facing report chart, not a raw debug plot and not a draft paper figure.
- Do not auto-send every training curve, per-step plot, or intermediate slice image.
- Preferred connector-chart palettes are Morandi-like and restrained:
sage-clay:#E7E1D6,#B7A99A,#7F8F84for the default QQ summary lookmist-stone:#F3EEE8,#D8D1C7,#8A9199for conservative summariesdust-rose:#F2E9E6,#D8C3BC,#B88C8Cfor secondary comparisons only
- Connector-facing chart requirements:
- white or near-white background
- low saturation, no neon colors
- one primary accent plus one neutral comparison color whenever possible
- simple legend, light grid, readable labels, and no dashboard clutter
- summarize only the evidence needed for the milestone
- Default chart choice:
- line chart for training / budget / step trends
- bar chart for a small number of categorical end-point comparisons
- point-range chart when uncertainty or seed spread matters
- If the figure encodes ordered magnitude, use a sequential muted palette; if it encodes signed delta around a reference, use a diverging muted palette with a neutral midpoint.
- Avoid rainbow / jet-like colormaps, 3D effects, and over-annotated dashboards.
- If the chart may later be reused in the paper, export a vector copy (
pdforsvg) alongside the connectorpng. - If the figure matters beyond transient debugging, open
figure-polish/SKILL.mdand follow its render-inspect-revise workflow before treating the image as final. - If plotting in Python, reuse the fixed Morandi plotting starter from the system prompt rather than inventing a new bright style for each run.
- If the runtime starts an auto-continue turn with no new user message, continue from the current run state, logs, artifacts, and active requirements instead of replaying the previous user turn.
- Progress message templates are references only. Adapt to the actual context and vary wording so messages feel human, respectful, and non-robotic.
- If a threaded user reply arrives, interpret it relative to the latest experiment progress update before assuming the task changed completely.
- Hard execution rule: every terminal command in this stage must go through
bash_exec; do not use any other terminal path for smoke tests, real runs, Git, Python, package-manager, or file-inspection commands. - Prefer
bash_execfor experiment commands so each run gets a durable session id, quest-local log folder, and laterread/list/killcontrol. - For meaningful long-running runs, include the estimated next reply time or next check-in window whenever it is defensible.
Tool discipline
- Do not use native
shell_command/command_executionin this skill. - All smoke tests, real runs, shell, CLI, Python, bash, node, git, npm, uv, and environment work must go through
bash_exec(...). - For git work inside the current quest repository or worktree, prefer
artifact.git(...)before raw shell git commands. - If a scratch repository or isolated test environment is needed, create and drive it through
bash_exec(...), not native shell tools.
Stage purpose
The experiment stage should turn a selected idea into auditable evidence. It should preserve the strongest old experiment-planning and execution discipline:
- define the run contract before execution
- keep the run comparable to baseline
- capture configs, commands, logs, and metrics
- report both success and failure honestly
- route the next action through an explicit decision
The experiment stage is not just "run code". It is the stage that converts an idea contract into evidence that other stages can trust. It is also the stage that should decide the next route once the measured result exists. Within the user's explicit constraints, maximize valid evidence per unit time and compute. Prefer equivalence-preserving efficiency upgrades first: larger safe batch size, mixed precision, gradient accumulation, dataloader workers, cache reuse, checkpoint resume, precomputed features, and smaller pilots. If a proposed efficiency change alters optimization dynamics, effective budget, or baseline comparability, treat it as a real experiment change and record it as such.
Use references/evidence-ladder.md when deciding whether the current package is merely executable, solid enough to carry the main claim, or already in the stage where broader polish is justified.
Completing one main run is not quest completion. After reporting the run, keep moving to iterate, analyze, write, or finalize unless a genuine blocking decision remains.
When the quest is algorithm-first, treat experiment as the execution surface of optimize, not as the terminal goal of the workflow.
After a measured result, the default next move is frontier review and optimize-side route selection rather than paper packaging.
Quick workflow
Treat this as the short run-order summary. The detailed run contract, execution rules, and recording rules remain in Workflow.
- Restate the selected idea in
1-2sentences and confirm the baseline comparison contract. - Before substantial code edits or the real main run, create
PLAN.mdandCHECKLIST.md. - Materialize or confirm a dedicated child
run/*branch/worktree for this main experiment line; one durable main experiment should map to one run branch and one Canvas node. - Use
PLAN.mdto lock the concrete run path, and useCHECKLIST.mdas the living control surface while planning, implementing, pilot testing, running, and validating. - Run a bounded smoke test or pilot before the real long run, then launch the real run with durable logging and monitor it through
bash_exec. - Once the route is concrete, prefer one clean implementation pass, one bounded smoke or pilot run, and then one normal main run; retry only after a concrete failure, invalidity, or genuinely new evidence justifies another attempt.
- Revise the plan if implementation, comparability, runtime, or route assumptions change materially, and close each real main-run milestone with a concise
1-2sentence summary that says what was tested, whether performance improved / worsened / stayed mixed, and the exact next action.
Non-negotiable rules
- Do not fabricate metrics, logs, claims, or improvement narratives.
- Do not introduce a new dataset or silently change splits or evaluation protocol.
- Do not change metric definitions or evaluation logic unless the change is explicitly justified and durably recorded.
- Do not stop after a quick sanity run if the agreed goal is a real experiment.
- Do not claim success before durable artifacts exist and the acceptance gate passes.
- Implement the claimed mechanism, not a convenient shortcut that changes the theory.
- Keep the baseline reference read-only.
- Avoid asking the user to fix the environment unless there is no credible agent-side path left.
- Do not record a durable main experiment from an idea branch, quest root branch, or paper branch as if that were the final result node; every durable main experiment should land on its own
run/*branch. - After each
artifact.record_main_experiment(...), route from the measured result:- if paper mode is enabled, decide whether to strengthen evidence, analyze, or write
- if paper mode is disabled, prefer iterate / revise-idea / branch over default writing
- In algorithm-first work, after each main run, return to
optimizeordecisionfor frontier review before launching another large run.
Experiment mental guardrails
- Baseline reproduction is not wasted time; untrusted comparison is wasted time.
- Failed runs are still data when the delta and diagnosis are recorded clearly.
- Suspiciously good results deserve the same skepticism as obvious failures.
- Change less, learn more.
- If a retry does not add new evidence, it is budget burn rather than progress.
Use when
- a baseline is accepted
- an idea has been selected
- the evaluation contract is explicit
- the quest is ready for implementation and measurement
Do not use when
- the baseline gate is unresolved
- the idea stage still has unresolved tradeoffs
- the main need is writing or follow-up analysis rather than a main run
Preconditions and gate
Before a main run starts, confirm:
- selected idea or hypothesis
- baseline reference
- dataset and split
- primary metric
- stop condition
- resource budget
- dedicated
run/*target branch or isolated worktree for this exact main experiment - exact output location
- required metric keys for acceptance
- minimal experiment and abandonment condition from the idea stage
If any of these are materially unknown, stop and resolve them through decision.
Required plan and checklist
Before substantial implementation work or a real main run, create a quest-visible PLAN.md and CHECKLIST.md.
- Use
references/main-experiment-plan-template.mdas the canonical structure forPLAN.md. - Use
references/main-experiment-checklist-template.mdas the canonical structure forCHECKLIST.md. PLAN.mdshould lead with the selected idea summarized in1-2sentences, put the user's explicit requirements and non-negotiable constraints first, and then make the run contract concrete: baseline and comparability rules, safe efficiency levers, code touchpoints, minimal code-change map, smoke / pilot path, full-run path, fallback options, monitoring and sleep rules, expected outputs, and a revision log.CHECKLIST.mdis the living execution list; update it during planning, implementation, smoke testing, main execution, validation, and every material route change.- If the code path, comparability contract, runtime strategy, or execution route changes materially, revise
PLAN.mdbefore spending more code or compute. - The later
RUN.md,summary.md, and artifact payloads remain required outputs, butPLAN.mdandCHECKLIST.mdare the canonical planning-and-control surface before and during execution. - Once
PLAN.mdmakes the implementation route concrete, do not keep reshaping code and commands speculatively. The normal default is one bounded smoke or pilot run and then one real run, with retries only after a documented failure, invalidity, or new evidence that changes the expected outcome.
Working-boundary rules
Only modify the active quest workspace for this experiment line.
- treat the accepted baseline workspace as read-only
- do not derive branch or worktree assumptions from guesswork
- keep all durable outputs inside the quest
- if the runtime gives an explicit worktree path, use it exactly
Resource and environment rules
- Follow the explicit resource assignment if one exists.
- If GPU assignment is explicit, respect it exactly and record it in the run manifest.
- Do not silently consume extra GPUs or broaden resource scope.
- Capture enough environment information that the run can later be reconstructed.
- If a new dependency appears necessary, record it as a risk and prefer a fallback if possible.
Truth sources
Use:
- idea-stage outputs
- baseline artifacts
- current codebase and configs
- recent decisions
- task and metric contract
- shell logs and generated outputs from the actual run
bash_execsession ids, progress markers, and exported logs from the actual run- the selected idea handoff contract
- incident or failure-pattern memory from earlier runs
Do not claim run success without durable outputs.
Required durable outputs
A meaningful experiment pass should leave behind:
- a run directory under
artifacts/experiment/<run_id>/or the quest-equivalent canonical location artifact_manifest.jsonrun_manifest.jsonmetrics.jsonmetrics.mdsummary.mdrunlog.summary.md- durable command, config, and log pointers
- exported shell log, typically
bash.log - a run artifact with explicit deltas versus baseline
- a decision about what should happen next
Recommended additional files:
claim_validation.md- environment snapshot files such as:
- Python version
- package freeze
- GPU info when applicable
- a live execution note or rolling run log when the experiment spans multiple implementation or execution steps
run_manifest.json should capture at least:
run_id- quest / branch context
- baseline reference or commit
- full commands
- config paths and key resolved hyperparameters
- dataset identifier or version
- seeds
- environment snapshot paths
- start time, end time, and final status
If a command needed for environment capture is unavailable, record that gap in the manifest and summary.
Workflow
1. Define the run contract
Before implementation or execution, state:
run_id- experiment tier:
auxiliary/devormain/test - research question
- null hypothesis
- alternative hypothesis
- hypothesis
- baseline id or variant
- metric targets
- expected changed files
- expected outputs
- stop condition
- compute or runtime budget
- minimal experiment
- abandonment condition
- strongest alternative hypothesis
- exact metric keys that will decide success or failure
Prefer to write this contract first in PLAN.md using references/main-experiment-plan-template.md, then keep the current execution state visible in CHECKLIST.md using references/main-experiment-checklist-template.md.
For substantial runs, also record the following seven experiment fields early and keep them updated during execution:
- research question
- research type
- research objective
- experimental setup
- experimental results
- experimental analysis
- experimental conclusions
If the run contract changes materially later, record the change durably.
Treat the run contract as a research question contract, not only an execution checklist. Before coding, be able to explain:
- why this run is the best current route rather than the main alternatives
- what observation would count as a real answer to the research question
- what result would force a downgrade, retry, or route change
- what confounder would make the run non-comparable even if it finishes successfully
If multiple candidate experiment packages exist, prefer the one with the best balance of:
- technical feasibility
- research importance
- methodological rigor
Do not choose a package only because it sounds ambitious.
For paper-facing lines, default to this evidence ladder:
auxiliary/dev- clarify parameters, settings, mechanisms, or diagnostics
main/test- carry the core comparison the paper will rely on
minimum -> solid -> maximum- first make the result executable and comparable
- then make it strong enough to carry the claim
- only then spend effort on broader supporting polish
2. Run a preflight check
Before editing or executing:
- confirm the dataset path, version, and split contract
- confirm the baseline metrics reference
- if durable state exposes
active_baseline_metric_contract_json, read that JSON file before planning commands or comparisons - treat
active_baseline_metric_contract_jsonas the default authoritative baseline comparison contract unless you record a concrete reason to override it - confirm the selected idea claim and code-level plan
- look up prior incidents or repeated failure patterns when available
- confirm output directories and naming
- confirm that the intended run still matches the current quest decision
If a repeated failure pattern already exists, apply the mitigation first and record that choice.
Also confirm before comparison work:
- the baseline verification is trustworthy enough
- the planned comparison still uses the same metric contract
- the metric keys and primary metric still match
active_baseline_metric_contract_jsonwhen that file is available - every main experiment submission still covers all required baseline metric ids from
active_baseline_metric_contract_json; extra metrics are allowed, but missing required metrics are not - the required baseline metrics still use the same evaluation code and metric definitions; if an extra evaluator is genuinely necessary, record it as supplementary output rather than replacing the canonical comparator
- if the run is
main/testand superiority is likely to be claimed, define the significance-testing plan before execution rather than after seeing the numbers - if
Result/metric.mdwas used during the run, treat it as optional scratch memory only and reconcile it against the final submitted metrics beforeartifact.record_main_experiment(...)
Before you begin a substantial run, send a concise threaded artifact.interact(kind='progress', ...) update naming:
- the run contract you are about to execute
- the main evidence it is testing
- the expected durable outputs
- the next checkpoint for reporting back
2.1 Diagnostic mode trigger
Switch from ordinary execution mode into diagnosis mode when any of the following becomes true:
- two retries in a row add no new evidence or no interpretable delta
- the baseline gap is much larger than expected and the cause is unclear
- the metrics are suspiciously strong, suspiciously identical to baseline, or highly unstable
- logs, checkpoints, or intermediate outputs conflict with the claimed behavior
In diagnosis mode:
- stop brute-force retrying
- prefer the smallest discriminative test that can separate competing hypotheses
- resolve obvious environment or data-contract issues before launching another comparison run
- make the diagnosis goal explicit: explain the behavior, not just "try something else"
3. Confirm the execution workspace
The normal experiment workspace is the current active idea worktree returned by artifact.submit_idea(...).
- do not create a fresh manual branch for the main experiment unless recovery or debugging truly requires it
- implement and run inside the current active idea workspace
- if the idea package changes materially before execution, submit a new durable idea branch with
artifact.submit_idea(mode='create', lineage_intent='continue_line'|'branch_alternative', ...)instead of silently mutating the old node - after a real main run finishes, record it with
artifact.record_main_experiment(...)before moving to analysis or writing - once that durable main result exists, treat the branch as a fixed round node; a later new optimization round should usually compare foundations and create a new
continue_linechild branch orbranch_alternativesibling-like branch - after
artifact.record_main_experiment(...), if QQ milestone media is enabled and the metrics are stable enough to summarize honestly, prefer one concise summary PNG over multiple attachments
4. Implement the minimum required change
Implementation rules:
- keep the change hypothesis-bound
- prefer small, explainable edits
- avoid unrelated cleanup during a main run
- record which files matter for later review
- preserve theory fidelity between the idea claim and the code change
- add robustness checks when the mechanism risks NaN, inf, or unstable behavior
- implement according to the current
PLAN.mdinstead of repeatedly improvising a new method after each small observation - avoid repeated code churn between the smoke test and the real run unless the smoke test exposes a specific problem that the next change is meant to fix
Prefer to complete one experiment cleanly before expanding to the next, unless parallel execution is explicitly justified and isolated. For substantial experiment packages, the default is one experiment at a time, with each one reaching a recoverable recorded state before the next begins.
Retry-delta discipline:
- unless the current state is completely non-executable, change only one major variable per retry
- if broader recovery is unavoidable, record exactly which layer changed: data, preprocessing, model, objective, optimization, evaluation, or environment
- before each retry, state the expected effect and the fastest falsification signal
- if the retry produced no interpretable delta, do not treat it as meaningful evidence about the underlying research hypothesis
5. Execute the run
Run with auditable commands and durable outputs.
Execution rules:
- use non-interactive commands
- prefer
bash_execinstead of ephemeral shell invocations - use the intended dataset and split
- keep logs durable
- report progress for long runs
- avoid silent metric-definition changes
- do not drift away from
active_baseline_metric_contract_jsonsilently when that file exists - avoid silently changing the baseline comparison recipe
- run the full agreed evaluation, not only a smoke test
You may do a quick sanity run first, but if the stage goal is a real experiment you must continue to the real evaluation unless the run is blocked and recorded.
Pilot-before-scale rule:
- start with a bounded pilot when the modification is non-trivial
- use the pilot to catch implementation mistakes early
- record pilot outcomes explicitly
- do not mistake pilot success for final evidence
Incremental-recording rule:
- do not wait until the end to reconstruct the run from memory
- update the durable run note after:
- contract definition
- important code changes
- pilot validation
- full execution checkpoints
- post-run analysis
- update
CHECKLIST.mdalongside those durable notes so the current execution frontier is obvious without replaying the whole log - include timestamps when they materially help reconstruction
- preserve failed attempts, anomalies, and partial outcomes rather than overwriting them
Last-known-good rule:
- keep track of the most recent state that was executable, comparable, and explainable
- when a new attempt breaks that state, debug forward from the last-known-good point instead of stacking more speculative edits on top of the broken state
- if the last-known-good state is unclear, reconstruct it before spending more budget on new hypotheses
5.1 Long-running command protocol
For commands that may run longer than a few minutes:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 1k
- Forks
- 119
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
ds-experiment- Source
- github.com/openlair/dr-claw