Experiment Controller

SkillAI & models

Use this skill whenever the user wants to execute experiments based on a finalized method and record results. Triggers include: 'run experiment', 'experiment controller', 'implement experiment', 'run model', 'execute training', 'experiment-controller', 'record results', 'ablation study', or any request to turn METHOD.md into concrete runs and output to EXPERIMENT.md. This skill is the **mandatory interface-layer experiment executor** in NeuroClaw: it searches literature/GitHub for matching experimental setups and codebases, proposes one scheme + repo after user discussion, uses git skills to download and setup, runs the experiment(s), and iteratively appends every result + observation to EXPERIMENT.md.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Experiment Controller skill

What this skill tells your AI

The instructions your AI receives, as published by cuhk-aim-group/neurodiscovery in skills/experiment-controller/SKILL.md and read by ahel’s review.

Overview

This skill implements the Literature/GitHub Search → Scheme Confirmation → Git Execution → Iterative Logging process for the NeuroClaw experiment-controller phase.

It acts as the Experiment Manager within the multi-agent framework:

  • Reads the latest IDEA.md and METHOD.md from the workspace.
  • Searches recent literature (multi-search-engine, arxiv-search, pubmed-search) and GitHub for reproducible experimental setups and open-source repositories that match the proposed architecture.
  • Summarizes candidate schemes (hyperparameters, datasets, baselines, training protocols) and proposes the most suitable GitHub repo.
  • Iteratively discusses with the user to confirm the exact scheme/repo.
  • After confirmation: uses git-essentials/git-workflows to clone, dependency-planner to install environment, claw-shell to run the experiment (training/inference/ablation).
  • After every run (or ablation), automatically records: setup details, metrics, logs, observations, and any issues.
  • Saves everything in EXPERIMENT.md (with dated sections for each run).

Research use only — the output is a complete, reproducible EXPERIMENT.md ready for paper-writing and future replication.

Quick Reference (Experiment Flow)

StepDescriptionOutput File
1. Read & ParseLoad IDEA.md + METHOD.md01_idea_method_summary.md
2. Literature & GitHub SearchFind matching setups & repos02_search_results.md
3. ProposalRecommend best scheme + repo03_proposal.md
4. User DiscussionConfirm scheme/repo04_discussion.md
5. Git Clone & SetupClone + install dependencies05_setup_log.md
6. Run ExperimentExecute training/inference/ablation06_run_log_*.md (per run)
7. Record ResultsAppend metrics + observationsEXPERIMENT.md

Harness Engineering Protocol (Mandatory for Reproducibility)

All experiments executed via experiment-controller must follow the Task Decomposition → Agent Initialization → Execution → Verification protocol to ensure self-validation, auditability, and resumability.

Phase 1: Task Decomposition

  • Parse experimental goal from METHOD.md into discrete tasks (data preprocessing, feature extraction, model training, inference, evaluation)
  • Define success criteria for each task (BIDS format compliance, feature vector shape/range, metric thresholds)
  • Identify cross-task dependencies and data flow
  • Output: experiment_task_manifest.json with task graph

Phase 2: Agent Initialization

  • Pre-flight checks: verify all dependencies installed, data paths accessible, Docker/GPU resources available (if required)
  • Environment snapshot: capture Python version, library versions, hardware specs, random seeds → environment_manifest.json
  • Checkpoint management: determine checkpoint frequency, rollback strategy, memory constraints
  • Logging setup: initialize structured logging with unique experiment session ID (timestamp + hash)

Phase 3: Execution with Incremental Logging

  • Execute each task with real-time progress tracking
  • Save checkpoint after each completed task (enables resumption on failure)
  • Log detailed metrics, intermediate outputs, and timing information
  • Generate cryptographic hash (SHA256) for each output artifact
  • Stream results to EXPERIMENT.md in real-time sections with timestamps

Phase 4: Verification and Post-Execution Validation

  • Self-verification checks (module-specific):
    • Preprocessing: verify BIDS compliance, check for NaN/Inf values, validate normalization ranges
    • Feature extraction: check output shape consistency, verify statistical properties (mean/std within expected bounds)
    • Model training: validate loss curve smoothness, check for NaN gradients, verify train/val split integrity
    • Inference: cross-check predictions for domain-specific constraints (probability bounds, anatomical plausibility)
  • Result integrity validation:
    • Recompute hash of all output files and compare with stored values
    • Flag any mismatches as potential corruption/tampering
  • Generate final audit report: experiment_audit_report.md with task execution times, success/failure status, verification results, and reproducibility metadata

Self-Verification Implementation Details

Each skill integrated into experiment-controller must include automatic validation steps:

Preprocessing validation:

- BIDS compliance: confirm file naming, JSON sidecars, required fields
- Data integrity: NaN/Inf count, range of pixel values, histogram sanity check
- Statistical bounds: mean/std within neuroimaging norms (e.g., T1w intensity ~0-4000 HU)

Feature extraction validation:

- Output shape check: row/column count match expected dataset size
- Distribution check: ensure features are not constant or degenerate
- Correlations: detect and warn on features with >0.95 mutual correlation

Training validation:

- Loss curve smoothness: flag sudden spikes or plateau too early
- Gradient health: ensure no NaN/Inf gradients during backprop
- Validation metric monotonicity (for early stopping): warn if validation improves inconsistently
- Train/val split: cross-verify split ratio and no subject leakage

Inference validation:

- Output shape consistency: predictions match expected target cardinality
- Domain constraints: probability predictions in [0,1], regression outputs within physiologically plausible ranges
- Batch effect check: compare results across batch sizes (should be near-identical with same seed)

Installation

# Place files in: skills/experiment-controller/

Important Notes & Limitations

  • Always starts from latest IDEA.md + METHOD.md; stops and prompts if missing.
  • Every step and every run is saved as numbered Markdown files for full transparency and resumption.
  • Git clone uses git-essentials/git-workflows (never manual commands outside skills).
  • Dependencies are handled exclusively by dependency-planner.
  • Multiple runs (e.g., ablations) are supported; results are appended with timestamps.
  • Final output always saved/updated as EXPERIMENT.md in workspace root.
  • Logs include: command executed, hyperparameters, metrics (accuracy, loss, Dice, etc.), runtime, observations, and any errors.

When to Call This Skill

  • Immediately after method-design completes METHOD.md
  • When the user wants to run, reproduce, or compare experiments
  • Before paper-writing (to populate quantitative results)

Complementary / Related Skills

  • research-idea → provides IDEA.md
  • method-design → provides METHOD.md (input)
  • multi-search-engine → literature & GitHub search
  • git-essentials / git-workflows → clone and manage repositories
  • dependency-planner → environment & dependency installation
  • claw-shell → internal execution of training scripts
  • paper-writing → consumes EXPERIMENT.md for results/tables

Reference

NeuroClaw architecture (section 1.7 experiment-controller skill). Flow: literature/GitHub search → user-confirmed scheme → git clone → dependency setup → iterative execution → EXPERIMENT.md logging.


Created At: 2026-03-24 00:00 HKT Last Updated At: 2026-04-05 02:01 HKT Author: chengwang96

Signals

GitHub stars
85
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
experiment-controller-cuhk-aim-group
Source
github.com/cuhk-aim-group/neurodiscovery