Evaluate ML Pipeline
SkillAI & modelsLets your agent run cross-validation scoring on a machine learning model and save the evaluation report.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Evaluate ML Pipeline skill
About this skill
Evaluate one learner with `skore.evaluate`. The splitter and any non-default score are already on the DataOp from `build-ml-pipeline`. This skill writes the call without `splitter=`, persists the report, and records the locator. When `model-ml-pipeline` dispatched this turn, return the locator. A st
What this skill tells your AI
The instructions your AI receives, as published by probabl-ai/skills in skills/evaluate-ml-pipeline/SKILL.md and read by ahel’s review.
Score one learner and persist the report. The pipeline, the
splitter, and any non-default score are already declared.
skore.evaluate is the entry point. Do not hand-roll
cross_val_score, cross_validate, classification_report,
or metric prints.
A SkrubLearner does not implement sklearn's fit(X, y).
cross_val_score raises. Call skore.evaluate(learner, data={...})
and omit splitter=.
Human-facing prose
Details: setup-workspace references/human_facing_prose.md.
Experiment markdown and # comments describe this evaluation.
Questions and the close use the same data-science language: not
skill ids, G-* names, or the wrapper CLI.
<!-- results-embed: … --> is a site marker.
Procedure
python -m skore_skills status. Unscaffolded workspace → setup. Thenpython -m skore_skills frame showwithout--revise. Anything other thanproceedloadsframe-ml-problemand stops.translationnull → stop. Do not write an evaluation call.- History-dependent pipeline (backward shift, lag, rolling
window, target shift, or a join with side history): if
tests/smoke/test_<stem>.pyis missing or pytest is red, stop. Route tobuild-ml-pipeline. Do not writeskore.evaluate. Documented n/a only when there is no history-dependent step. python -m skore_skills evaluate consent --stem <stem>.ask, and this turn has not already answered Evaluate frombuild-ml-pipeline: quote JSONcontextand present Evaluate (Recommended) / Modify / Stop. Stop. "Run evaluation" is not consent on a first run.- The user already answered Evaluate this turn, or JSON
is
proceed(a real locator already exists): continue. stop— smoke file missing. Route to build.
- G-SKORE-MODE. Read
status.policy.skore_mode. If it is already set, keep it. If unset, ask local (recommended) / hub / mlflow, thenpython -m skore_skills policy set skore_mode <mode>. local:mkdir reports(exist_ok); do not writereports/README.md. hub or mlflow: do not createreports/. Then loadadd-python-packageif installed; it runsenv add-skore --mode <mode> --execute. Do not spellskore[...]here. Do not runenv addhere. Constructors:references/g_skore_mode.md. A switch or a migration loadssync-ml-reportswhen that skill is installed. - Emit Pre-flight, then 1–3 sentences: this is local
full-dataset evaluation with the locked scheme, report
metrics, and
project.put. Name the stem, the fold count when known, andexperiments/<stem>.pyplusscratch/results/<stem>/. Timing depends on rows, folds, and the learner. Do not invent minutes. If consent is still pending, preview and stop. - Write the call in
experiments/NN_*.pyonly. See the call shapes below.python -m skore_skills styleafter the edit. - After
put, record the locator, write the snapshots, then End of turn.
Every Python probe goes to scratch/<ts>_<short>.py and runs
with the composed-dev Python from env verify. No inline
python -c. No warnings.filterwarnings unless the user asks.
An unconfirmed signature stays unwritten. End that probe with
BLOCKED: <class> signature needs an API lookup that cannot run this turn (<why>). A cache hit is a satisfied lookup.
The call
skore.evaluate in experiments/NN_*.py. Omit splitter= so
skore reuses the DataOp cv and split_kwargs. Passing
splitter= drops split_kwargs. Omitted splitter= with no
DataOp cv is an 80/20 holdout, correct only when
translation.report is EstimatorReport. Wiring:
references/metadata-routing.md.
SkrubLearner—skore.evaluate(learner, data={...}). Keys are theskrub.varnames. Interop:references/skrub_interop.md.- An estimator whose
fitis(X, y)—skore.evaluate(estimator, X, y). Still omitsplitter=when the lockedcvis already the evaluation scheme.
The cv on mark_as_X is KFold, GroupKFold, or the
date-based class from build. If translation names a cv and
the marker has none, return to build-ml-pipeline. Do not
wire split_kwargs here. Empty split_kwargs plus a possible
group column → return to build-ml-pipeline. Do not default
to KFold.
No Stratified* for class imbalance. It compresses across-fold
variance.
The headline is translation.metric. If build attached
.skb.with_scoring(...), inspect report.metrics.score().
skore.evaluate has no scoring= argument. A metric or a
sample_weight that belongs on the DataOp goes back to build
(references/custom-metrics.md).
An extra check the user asks for after the lock:
references/custom-checks.md. Subclass skore.Check at module
level in experiments/NN_*.py, then report.checks.add(...)
after evaluate and before project.put. add extends SKD
checks. Do not invent a check. Do not register one from
audit/. Confirm Check and checks.add with api get.
Escalate past evaluate only when the dispatcher is too coarse
(references/reports.md): EstimatorReport for one held-out
fit, CrossValidationReport for per-fold artifacts,
ComparisonReport for two or more learners. Holdout uses
EstimatorReport.
CV is necessary but not sufficient for any pipeline with
history-dependent features. skore.evaluate materializes the
graph once with one env-dict and splits indices. The smoke test
exercises a fresh env-dict at predict time. A passing smoke
test is still required before the caller may flip status to
done. Do not edit JOURNAL.md History or the design-note
Status to done from this skill.
skore.evaluate(...) and project.put(...) live only in
experiments/NN_*.py. A scratch probe, an audit file, or a
notebook that calls them duplicates the report under the same
key. Read a stored report with project.summarize() then
project.get(id). get(key) raises KeyError because get
is by id. Do not re-run evaluate to paper over that.
Plots that skore does not already draw load plot-ml-figure
when that skill is installed.
After put
Project.put returns None. Read
project.summarize().frame(), take the newest row for the key,
and form one locator. Write that exact string to
scratch/results/<stem>/locator.txt. Run
python -m skore_skills loop locator --stem <stem> and paste
JSON locator verbatim. Missing locator is
n/a — backend did not expose a locator.
- local —
local workspace: [reports/](../reports/) · id: <id>, plus the absolutereports/path. Do not link a private file. - hub — the exact
Consult your report at …URL:[Open report](<url>) · hub · id: <id>. If Skore emits no URL, link the project landing page asOpen project. - mlflow — an exact
View run …URL when present. If absent andtracking_uriis HTTP(S), use/#/experiments/<experiment-id>/runs/<run-id>. Forfile:,sqlite:,databricks, or any other URI without an emitted URL:mlflow · tracking: <uri> · experiment: <id> · run: <run-id>. Do not invent a browser link.
In the same experiment file, write report._repr_html_() to
scratch/results/<stem>/report.html and repr(report) to
report.txt. Overwrite pipeline.html from a fitted
estimator: report.estimator_ on EstimatorReport,
report.reports_[0].estimator_ on CrossValidationReport.
Prefer _repr_html_ when the fitted object defines it.
Confirm with api get. Do not call SkrubLearner.report or
full_report.
End of turn
When model-ml-pipeline dispatched this turn, pass JSON
locator up and return. Do not run loop artifacts, audit,
record-outcome, convert, site, or git end-turn.
Otherwise this skill owns the close. When both policy.notebooks
and policy.site are true, the turn is unfinished until
record-outcome, notebook convert --html, site build, and
git end-turn have run, in that order, after the locator is in
the close. A stop at an earlier gate still names notebook convert
then site build in that order. Do not leave them as a conditional
aside, and do not stop after the narrative. Run
python -m skore_skills loop artifacts --stem <stem>.
stop / evaluate_incomplete → name the missing file and do
not audit. audit → load audit-ml-pipeline when installed.
record → skip audit.
Then the checkpoint. If audit ran, wait for its close and pass
its digest and G-AUDIT-FINDING into manage-ml-backlog
record-outcome when that skill is installed. If audit did not
run, call record-outcome with the user's headline, if any, and
n/a — audit not run. Missing backlog skill → one line. Do
not write History here. Never mark done while smoke is red.
Record-outcome runs before convert and site build.
If policy.notebooks is true and export-ml-notebook is
installed, run python -m skore_skills notebook convert experiments/<stem>.py, with --html when policy.site is
also true. Convert re-executes the script. Missing converters
→ one line naming add-python-package. Then, if
policy.site is true and export-ml-site is installed, run
python -m skore_skills site build. A build error does not
fail the turn.
Run python -m skore_skills git end-turn --stage evaluate.
If JSON action is invoke, load persist-ml-git when
installed and stop. Otherwise load triage-ml-task when
installed. Do not run git commit here.
Checkpoint
Standalone only. 2–6 sentences of the result, grounded in the
headline or, when audit was skipped, report.txt. If audit
ran, ground the story in its Checks and Metrics. Do not invent
a metric.
Links: when site build ran or is about to,
[report.html](<workspace>/report.html) and
html/<stem>.html. Otherwise
[journal/<stem>.md](journal/<stem>.md).
Tokens after the narrative: JSON locator first, then
G-AUDIT-FINDING (n/a — audit not run when skipped).
Stops
- No scaffold (
has_srcandhas_journalboth false) → setup. Do not requiregit. - Smoke missing or red on a history-dependent pipeline → build.
- Empty
split_kwargsplus a possible group column → return tobuild-ml-pipeline. Do not default toKFold. import skorefails → G-SKORE-MODE if unset, thenadd-python-package. Do not drop back tocross_val_score.- Hyperparameter search, serving, and multi-run tracking are out of scope.
- A missing package loads
add-python-packagewhen installed. Do not runenv addhere.
Pre-flight — emit before any code
Pre-flight (evaluate-ml-pipeline):
- [ ] sklearn, skrub, skore import
- [ ] frame show is proceed; DataOp cv matches translation
(or holdout, and the marker has no cv)
- [ ] skore_mode is set (local | hub | mlflow)
- [ ] evaluate consent is proceed, or the user answered Evaluate
- [ ] Call site is experiments/NN_*.py
- [ ] skore.evaluate omits splitter=
- [ ] Smoke: passing | n/a (no history-dependent step) | STOP
References
references/metadata-routing.md— the lockedcvstays on the DataOp; evaluate omitssplitter=.references/skrub_interop.md— env-dict versus(X, y).references/g_skore_mode.md— Project constructors.references/reports.md— whenevaluateis too coarse.references/custom-metrics.md— metric kwargs andwith_scoring.references/custom-checks.md— a check the user asked for.
Signals
- GitHub stars
- 132
- Forks
- 9
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
evaluate-ml-pipeline- Source
- github.com/probabl-ai/skills