Experiment Controller
SkillAI & modelsUse this skill whenever the user wants to execute experiments based on a finalized method and record results. Triggers include: 'run experiment', 'experiment controller', 'implement experiment', 'run model', 'execute training', 'experiment-controller', 'record results', 'ablation study', or any request to turn METHOD.md into concrete runs and output to EXPERIMENT.md. This skill is the **mandatory interface-layer experiment executor** in NeuroClaw: it searches literature/GitHub for matching experimental setups and codebases, proposes one scheme + repo after user discussion, uses git skills to download and setup, runs the experiment(s), and iteratively appends every result + observation to EXPERIMENT.md.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Experiment Controller skill
What this skill tells your AI
The instructions your AI receives, as published by cuhk-aim-group/neuroclaw in skills/experiment-controller/SKILL.md and read by ahel’s review.
Overview
This skill implements the Literature/GitHub Search → Scheme Confirmation → Git Execution → Iterative Logging process for the NeuroClaw experiment-controller phase.
It acts as the Experiment Manager within the multi-agent framework:
- Reads the latest IDEA.md and METHOD.md from the workspace.
- Searches recent literature (multi-search-engine, arxiv-search, pubmed-search) and GitHub for reproducible experimental setups and open-source repositories that match the proposed architecture.
- Summarizes candidate schemes (hyperparameters, datasets, baselines, training protocols) and proposes the most suitable GitHub repo.
- Iteratively discusses with the user to confirm the exact scheme/repo.
- After confirmation: uses git-essentials/git-workflows to clone, dependency-planner to install environment, claw-shell to run the experiment (training/inference/ablation).
- After every run (or ablation), automatically records: setup details, metrics, logs, observations, and any issues.
- Saves everything in EXPERIMENT.md (with dated sections for each run).
Research use only — the output is a complete, reproducible EXPERIMENT.md ready for paper-writing and future replication.
Quick Reference (Experiment Flow)
| Step | Description | Output File |
|---|---|---|
| 1. Read & Parse | Load IDEA.md + METHOD.md | 01_idea_method_summary.md |
| 2. Literature & GitHub Search | Find matching setups & repos | 02_search_results.md |
| 3. Proposal | Recommend best scheme + repo | 03_proposal.md |
| 4. User Discussion | Confirm scheme/repo | 04_discussion.md |
| 5. Git Clone & Setup | Clone + install dependencies | 05_setup_log.md |
| 6. Run Experiment | Execute training/inference/ablation | 06_run_log_*.md (per run) |
| 7. Record Results | Append metrics + observations | EXPERIMENT.md |
Harness Engineering Protocol (Mandatory for Reproducibility)
All experiments executed via experiment-controller must follow the Task Decomposition → Agent Initialization → Execution → Verification protocol to ensure self-validation, auditability, and resumability.
Phase 1: Task Decomposition
- Parse experimental goal from METHOD.md into discrete tasks (data preprocessing, feature extraction, model training, inference, evaluation)
- Define success criteria for each task (BIDS format compliance, feature vector shape/range, metric thresholds)
- Identify cross-task dependencies and data flow
- Output:
experiment_task_manifest.jsonwith task graph
Phase 2: Agent Initialization
- Pre-flight checks: verify all dependencies installed, data paths accessible, Docker/GPU resources available (if required)
- Environment snapshot: capture Python version, library versions, hardware specs, random seeds →
environment_manifest.json - Checkpoint management: determine checkpoint frequency, rollback strategy, memory constraints
- Logging setup: initialize structured logging with unique experiment session ID (timestamp + hash)
Phase 3: Execution with Incremental Logging
- Execute each task with real-time progress tracking
- Save checkpoint after each completed task (enables resumption on failure)
- Log detailed metrics, intermediate outputs, and timing information
- Generate cryptographic hash (SHA256) for each output artifact
- Stream results to EXPERIMENT.md in real-time sections with timestamps
Phase 4: Verification and Post-Execution Validation
- Self-verification checks (module-specific):
- Preprocessing: verify BIDS compliance, check for NaN/Inf values, validate normalization ranges
- Feature extraction: check output shape consistency, verify statistical properties (mean/std within expected bounds)
- Model training: validate loss curve smoothness, check for NaN gradients, verify train/val split integrity
- Inference: cross-check predictions for domain-specific constraints (probability bounds, anatomical plausibility)
- Result integrity validation:
- Recompute hash of all output files and compare with stored values
- Flag any mismatches as potential corruption/tampering
- Generate final audit report:
experiment_audit_report.mdwith task execution times, success/failure status, verification results, and reproducibility metadata
Self-Verification Implementation Details
Each skill integrated into experiment-controller must include automatic validation steps:
Preprocessing validation:
- BIDS compliance: confirm file naming, JSON sidecars, required fields
- Data integrity: NaN/Inf count, range of pixel values, histogram sanity check
- Statistical bounds: mean/std within neuroimaging norms (e.g., T1w intensity ~0-4000 HU)
Feature extraction validation:
- Output shape check: row/column count match expected dataset size
- Distribution check: ensure features are not constant or degenerate
- Correlations: detect and warn on features with >0.95 mutual correlation
Training validation:
- Loss curve smoothness: flag sudden spikes or plateau too early
- Gradient health: ensure no NaN/Inf gradients during backprop
- Validation metric monotonicity (for early stopping): warn if validation improves inconsistently
- Train/val split: cross-verify split ratio and no subject leakage
Inference validation:
- Output shape consistency: predictions match expected target cardinality
- Domain constraints: probability predictions in [0,1], regression outputs within physiologically plausible ranges
- Batch effect check: compare results across batch sizes (should be near-identical with same seed)
Installation
# Place files in: skills/experiment-controller/
Important Notes & Limitations
- Always starts from latest IDEA.md + METHOD.md; stops and prompts if missing.
- Every step and every run is saved as numbered Markdown files for full transparency and resumption.
- Git clone uses git-essentials/git-workflows (never manual commands outside skills).
- Dependencies are handled exclusively by dependency-planner.
- Multiple runs (e.g., ablations) are supported; results are appended with timestamps.
- Final output always saved/updated as
EXPERIMENT.mdin workspace root. - Logs include: command executed, hyperparameters, metrics (accuracy, loss, Dice, etc.), runtime, observations, and any errors.
When to Call This Skill
- Immediately after method-design completes METHOD.md
- When the user wants to run, reproduce, or compare experiments
- Before paper-writing (to populate quantitative results)
Complementary / Related Skills
research-idea→ provides IDEA.mdmethod-design→ provides METHOD.md (input)multi-search-engine→ literature & GitHub searchgit-essentials/git-workflows→ clone and manage repositoriesdependency-planner→ environment & dependency installationclaw-shell→ internal execution of training scriptspaper-writing→ consumes EXPERIMENT.md for results/tables
Reference
NeuroClaw architecture (section 1.7 experiment-controller skill). Flow: literature/GitHub search → user-confirmed scheme → git clone → dependency setup → iterative execution → EXPERIMENT.md logging.
Created At: 2026-03-24 00:00 HKT Last Updated At: 2026-04-05 02:01 HKT Author: chengwang96
Signals
- GitHub stars
- 85
- Forks
- 4
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
experiment-controller- Source
- github.com/cuhk-aim-group/neuroclaw