Research ML Practice
SkillDev toolsLets your agent work as a research Claude skill that looks up best practices for ML questions like leakage, transforms, and feature choice.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Research ML Practice skill
About this skill
Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus existing EDA. Trigger when explore-ml-data or build-ml-pipeline load this skill. Not for routine profiling o
What this skill tells your AI
The instructions your AI receives, as published by probabl-ai/skills in skills/research-ml-practice/SKILL.md and read by ahel’s review.
Worker skill. Callers own the stage turn and the user-facing
summary. Do not git end-turn.
Human-facing prose
Details: setup-workspace references/human_facing_prose.md.
Chat stays path + one or two sentences of the finding. Do not
narrate the skills framework or the wrapper CLI.
Sequence
- Intake. Infer modality from dtypes / JOURNAL when
obvious; use a stated domain and stage (
data_analysis|model) if the caller named them. Then pick a mode (references/search.md):- Caller passed a named concern (not the canned EDA extra-analysis question) → depth. Do not re-ask. Do not run the canned survey first.
- Caller passed the canned extra-analysis survey or “survey extra analyses” → survey.
- Neither → if JOURNAL and an EDA report exist, run the canned survey; else AskUserQuestion for a named concern, stating the modality and stage inferred this turn and that the answer only drives a literature search, no code. A file link is an addition, never the context. Do not start a depth search with an empty concern.
- Abstract the problem class before any query
(
references/search.md). JOURNAL and EDA are context, not the answer list and not search keywords for the table’s proper name. - Run the matching search loop (
references/search.md). Fetch primary pages. Follow up per promising extra if the first pass is thin or single-sourced. - Write
scratch/research/using the matching structure inreferences/search.md(survey:survey-<slug>.md; depth:<slug>.mdwith lanes). There is notemplates/directory in this skill — copy the markdown skeleton from that reference. Gitignored. - Return to the caller: scratch path and a one- or two-sentence
finding. Name that candidates are laned (
measure/declare/evaluate/confirm). Chat is path + those sentences only — no pasted headings, tables, or “Depth note — …” body. That is the complete user-facing deliverable even when a harness says to put the full answer in chat. If tools cannot search or write, stop there: still only path + sentences (name the intendedscratch/research/path). No hypotheses, diagnostics, planned-query bullets, or template headings in chat (that is the paste). Do not writedata_analysis.md, the design note, ordata/. The caller asks which extras to add.
Stop conditions
- Do not drop, impute, or remove outliers. Do not change the split.
- Do not pick the final learner or architecture.
- Do not treat a single blog as ground truth; say when sources disagree. Thresholds need two independent sources.
- Do not
pixi add/uv add/env add. If code needs a library, nameadd-python-packageand return. - Do not run
api getas a substitute for literature (symbols still go throughapi getin the caller). - Missing skill from a caller → that caller one-line skips.
- Do not copy Open questions / EDA findings onto the extras list without a source.
- Do not search the dataset proper name,
sklearn.datasets, a Kaggle slug, or “baseline pipeline”. Do not return learners /Pipelinesteps as EDA extras. - Never answer from memory when search ran. Do not ask the user to go look something up.
- Do not paste the scratch markdown into chat (no “Depth note —”, no survey body). Path + 1–2 sentences only. If tools cannot search or write, stop after that. A no-tools harness does not license pasting the note.
Signals
- GitHub stars
- 132
- Forks
- 9
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
research-ml-practice- Source
- github.com/probabl-ai/skills