Review ML Experiment
SkillFiles & storageThis review Claude skill audits a finished ML experiment after your approval and writes one idea file per improvement candidate.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Review ML Experiment skill
About this skill
Post-evaluate review. Gate the expensive skore-check audit, then write one markdown idea file per candidate. Trigger after a successful evaluate, on "review this stem", or when review consent is Review or proceed. Do not write JOURNAL.md or a design note. Do not run cells run until the user accepts
What this skill tells your AI
The instructions your AI receives, as published by probabl-ai/skills in skills/review-ml-experiment/SKILL.md and read by ahel’s review.
Optional loop step after evaluate. Record-outcome stays with the caller. This skill writes idea files only.
Human-facing prose
Details: setup-workspace references/human_facing_prose.md.
The Review / Skip / Stop question and cost preview describe
reading this report and writing follow-up ideas — not skill ids,
cells run, or the wrapper CLI.
Procedure
- Run
python -m skore_skills statusandpython -m skore_skills review consent --stem <stem>. Treat JSONactionas authoritative.stop— noscratch/results/<stem>/report.html. Name that file and stop. Do not audit.ask— report exists, digest does not. Emit the cost preview below, then AskUserQuestion: Review (Recommended) / Skip / Stop. "Audit it" on a first run is not consent. If this turn already answered Review, do not ask again.proceed— digest already on disk. Do not re-run checks. Refresh idea files from the existing digest.
- Skip — write no idea files. Return
n/a — audit not runso the caller can record-outcome. - Stop — do not audit and do not record-outcome. The persisted
report stays. Load
triage-ml-taskwhenstatus.skills.triage-ml-taskis true. - Review, or
proceedfor idea refresh: loadaudit-ml-pipelineonly ifstatus.skills.audit-ml-pipelineis true. On Review, that skill runscells run(consent staysaskuntil the digest exists; the Review answer is what starts it). Onproceed, do notcells run. Missing audit skill → one-line skip; returnn/a — audit not runand write no idea files. Do not open the Project or callreport.*here. - Read the design note, the EDA summary, the last History row,
and the digest. One candidate per
Issues:/Tips:line. A methodological gap the design note named and this run did not test is another candidate. A user idea or a literature query is not a candidate here: after this skill returns,shape-user-ideaorsearch-ml-literaturewrites that file when the user asks. Loadresearch-ml-practiceonly ifstatus.skills.research-ml-practiceis true and an audit or design candidate needs sources; otherwise one-line skip. Do not invent papers, metrics, or a winner. - Write one file per candidate at
journal/ideas/<stem>-<slug>.mdwith Experiment, Source (audit:<stem>:checks.<code>ordesign:<stem>), Triageopen, Question, Why now, What changes, Open gaps. No acceptance criteria. On a refresh, keep an existing file'sTriagevalue. A new candidate isopen. - Return the digest, JSON
findingfrompython -m skore_skills audit finding --stem <stem>, the locator frompython -m skore_skills loop locator --stem <stem>, and the idea paths.
An explicit re-audit asks the gate again before cells run, even
when a digest exists, because it re-runs the checks.
Cost preview
On ask, before the question, emit 1–3 sentences. This is a
local read of the persisted report, not another fit. The audit
template runs every skore check, writes audit/<stem>.py and
scratch/audit/<stem>/audit.md, and can be slow on a large
report. Name the stem and those paths. Do not invent minutes.
Stop conditions
- Do not run
cells runbefore Review or an explicit re-audit confirmation. - Do not write
JOURNAL.mdor a design note. - Do not call
skore.evaluateorproject.put. - Do not pick a winning idea.
- Do not invent a missing child's procedure.
Signals
- GitHub stars
- 132
- Forks
- 9
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
review-ml-experiment- Source
- github.com/probabl-ai/skills