Build ML Pipeline

SkillDev tools

Lets your agent build ML pipelines as skrub DataOps graphs when you ask it to build Claude skill projects.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Build ML Pipeline skill

About this skill

Declare the pipeline from data source to predictor as a skrub DataOps graph. Stateless steps use `.skb.apply_func`. Stateful steps use `.skb.apply`. After the graph exists, attach the locked splitter and any non-default score. No fit, tune, persistence, or `skore.evaluate`.

What this skill tells your AI

The instructions your AI receives, as published by probabl-ai/skills in skills/build-ml-pipeline/SKILL.md and read by ahel’s review.

Declare a skrub DataOps graph, then attach the locked splitter and any non-default score. Smoke is the last step here. Do not fit, tune, persist, or call skore.evaluate.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Experiment markdown, design-note Method text, and # comments describe this pipeline. Questions and the checkpoint use the same data-science language: not skill ids, G-* names, or the wrapper CLI. <!-- results-embed: … --> is a site marker. Authoring hints stay in this skill.

X marker is .skb.mark_as_X(). Predict grid is the rows to score. Cross-row step reads other rows (lag, rolling, group-agg, side join). Layers 1 / 2 / 3 are sources, the grid plus marker, and features after the marker.

Procedure

  1. python -m skore_skills status. Missing scaffold → S0. Then python -m skore_skills design consent --stem <stem>. JSON action is authoritative. ask / stop → do not declare ("build it" is not approval). proceed continues. Then python -m skore_skills frame show without --revise. Anything other than proceed loads frame-ml-problem and stops. A proceed whose translation is null stops: there is no splitter translation.
  2. Emit 1–3 sentences before the first write: this is local preparation of an unfitted build_learner, Method cells, and a pipeline snapshot, then smoke. It does not train or run full-dataset evaluation. Do not invent a duration. Say it once. If approval is pending, preview only.
  3. Emit Pre-flight. Tick a box only with evidence from this turn.
  4. Declare build_learner (Rule 1–3). Sources, features, and the estimator first. Then attach cv and any non-default score from translation, still before .skb.make_learner(). python -m skore_skills api get for every new symbol. python -m skore_skills style after edits. Probes go to scratch/ via the composed-dev Python from env verify. No inline python -c. No warning filters unless the user asks.
  5. Unfitted snapshot: references/snapshot.md.
  6. When experiments/NN_*.py exists for this stem, load smoke-test-ml-pipeline only if that skill is installed. Missing skill → one line; do not invent the pytest file. Then python -m skore_skills smoke run --stem <stem>. JSON stop / red / smoke_missing → fix the graph here. Do not loosen the assertion. Do not evaluate. Do not claim pytest is green without this command.
  7. On smoke run proceed, the checkpoint below, then python -m skore_skills evaluate consent --stem <stem>.
    • ask — quote JSON context (question, experiment, smoke, persisted_report) in 2–4 lines: what the answer authorizes, the design question, and that this stem has no persisted report yet. Then AskUserQuestion, Evaluate (Recommended) / Modify / Stop, and stop. Do not write skore.evaluate. Modify → edit, then smoke run again. Stop → end.
    • proceed — a report already exists. Load evaluate-ml-pipeline if installed, or return to model-ml-pipeline when that skill called this one.
    • stop — smoke file missing. Do not evaluate. If the user answers Evaluate, load evaluate-ml-pipeline when installed. smoke run proceed is not that answer.

Checkpoint

After smoke run proceed, before the Evaluate menu. 2–6 sentences: what was declared, that smoke is green, the learner. Ground them in Method. Do not invent a metric. No locator yet.

Then links. When site build ran this turn: [report.html](<workspace>/report.html) and html/<stem>.html. Otherwise [journal/<stem>.md](journal/<stem>.md) and [experiments/<stem>.py](experiments/<stem>.py).

Entry contracts

Implement the approved design. Do not silently upgrade it.

  • Locked baseline — each token in the journal baseline cell, in order. dummy is DummyClassifier / DummyRegressor in the normal DataOps graph; it checks that the path runs. logistic, seasonal_naive, group_mean, and production are comparison models. api get the class. Several tokens are several comparison models.
  • EDA-backed — only Method-cited findings. A missing choice stops for a question. A temporal finding is the locked cv, not a license for three layers, lags, or AlignXy unless Method names those steps.
  • Backlog / discussion — Method is the boundary.

On a feature, transform, or leakage question, load research-ml-practice if it is installed. Abstract the problem class. AskUserQuestion allow_multiple on declare rows that do not violate a stop. measure revisits EDA and does not edit data_analysis.py. The splitter still comes from frame show. evaluate names evaluate-ml-pipeline. confirm asks the user. Missing skill → one line.

Rule 1 — Skrub DataOps is the pipeline entry point

Root at skrub.var(...), not a bare sklearn.Pipeline. skrub.X / skrub.y are not roots (S4). If the user asks for sklearn.Pipeline or build_pipeline(), do not import Pipeline, even as an inner estimator. Redirect to skrub.var and build_learner returning predictions.skb.make_learner(). Do not illustrate the refusal with a Pipeline([...]) constructor.

import skrub
from sklearn.ensemble import HistGradientBoostingRegressor

from <pkg>.data import TARGET_COL, load_raw


def build_learner(data_dir_preview=None):
    """Return the unfit learner (skrub SkrubLearner)."""
    data_dir = (
        skrub.var("data_dir", value=str(data_dir_preview))
        if data_dir_preview is not None
        else skrub.var("data_dir")
    )
    data = data_dir.skb.apply_func(load_raw)
    X = data.drop(columns=[TARGET_COL]).skb.mark_as_X()
    y = data[TARGET_COL].skb.mark_as_y()
    predictions = X.skb.apply(
        HistGradientBoostingRegressor(random_state=0), y=y
    )
    return predictions.skb.make_learner()

A quick traditional baseline uses skrub.tabular_pipeline (or TableVectorizer) plus one task-appropriate estimator. Confirm both with api get. Do not hand-tune columns or search hyperparameters here. Other shapes: references/common_patterns.md.

Rule 2 — Mark X early; featurize after

Per-row math and fit-time encoders: marker on the loaded frame. Any cross-row step: marker upstream of that step. Code for the three layers, including the loader-baked horizon refusal: references/layer_examples.md. Read it before proposing Layer 2.

value= is preview only. Expose data_dir_preview=None on build_learner. Do not bake a relative path into pipeline.py.

Splitter and scoring, after the graph

Read translation from the frame show that returned proceed. Attach on the existing X marker, then .skb.make_learner(). skrub requires cv= whenever split_kwargs is set. Integer cv is not a splitter. Evaluate omits splitter= so skore reuses this cv (evaluate-ml-pipeline/references/metadata-routing.md).

That cv copies the locked deployment: a scored fit contains only rows that deployment would already have seen. The shapes below are the usual ones. When the deployment is a different structure, open references/custom-splitter.md and write split for that setting.

  • scheme date_time — open the time series section of references/custom-splitter.md. cv= is the project-local class that section describes. Build one of those splitters per entry in translation.horizons. Each one is one predictor. split_kwargs carries the timestamp values. The timestamp column is the one the EDA or the text shipped with the data already names. Ask which column holds the timestamps only when those sources do not name one. time_role covariate keeps that column in the features. sort_key drops it from the features and still passes it in split_kwargs.
  • splitter GroupKFold — cv=GroupKFold(n_splits=<folds>) and split_kwargs={"groups": data["<translation.groups>"]}.
  • splitter KFold — cv=KFold(n_splits=<folds>) and empty split_kwargs.
  • Holdout (report EstimatorReport) or translation null — no cv.

Scoring uses translation.metric. If that name is one skore already reports for the task (regression: MSE, RMSE, MAE, R²; binary: accuracy, precision, recall, F1, ROC-AUC; multiclass: macro and micro variants; multilabel: per-label and averages), do not attach a scorer. Otherwise .skb.with_scoring(...) in this same late step. The callable shape is evaluate-ml-pipeline/references/custom-metrics.md. Do not pass scoring= to skore.evaluate.

Do not write skore.evaluate(...) here. Do not call train_test_split from pipeline code.

Rule 3 — Attach: stateless function, stateful estimator

  • .skb.apply_func(fn) — output depends only on the current row and constants.
  • .skb.apply(estimator) — learns on training, reapplies on test.
  • skrub.deferred — rare; only when combining several DataOps and no skrub joiner fits. Default is apply_func. Details: references/source-binding.md.

Would the output change on the training subset versus the whole frame? Yes → .skb.apply. Means, medians, quantiles, vocabularies, target encoding, TF-IDF: stateful.

STOP — target encoding / apply_func. When the user asks for
`def target_encode` + `.skb.apply_func`: refuse. Do not paste
the leaky function body and then the fix. Propose sklearn
TargetEncoder (or BaseEstimator + TransformerMixin) via
`.skb.apply`. Name `api get` for the signature.

Reproducibility

done History rows stay runnable. Details: references/reproducibility_mechanics.md.

  • Option 1 — a default-preserving flag on the existing function. Small append. The default keeps prior callers unchanged (include_calendar_features: bool = False).
  • Option 2 — a new function called only from the new experiment.
  • Option 3 — branch the module. Last resort.

Three or more flags, or a flag that changes an existing caller's default → Option 2, or stop. After the change, pytest all of tests/smoke/.

Stops

S0. Workspace not scaffolded

status first. has_src and has_journal both false → stop. Do not require git. Missing design: design consent; ask / stop stay here. Missing data contract: explain and stop.

S0b. Modeling decisions are not locked

  • Rule: status.modeling_decisions must be locked, and python -m skore_skills frame show must return proceed with a non-null translation, before any declaration.
  • Recovery: load frame-ml-problem when installed. A null translation stops with no pipeline. Do not invent the table.

S1. Missing dependency

import skrub / sklearn failure, or a DataOp HTML stub, loads add-python-package for skrub and scikit-learn. Do not env add here. Do not substitute sklearn.Pipeline.

S2. Symbol from memory is forbidden

Every new skrub, sklearn, or skore name comes from api get or a matching cache read this turn.

S4. skrub.X / skrub.y are not graph roots

Root on skrub.var("<source>", value=preview). An existing skrub.X graph: show the alternative and ask. Do not auto-rewrite. Catalogue: references/source-binding.md.

Refuse: skrub.X / skrub.y are not graph roots (S4).
They bake the marker at the source and defeat Layer 1.

Alternative (refactor — ask before rewriting):
  data = skrub.var("data_dir", value=preview).skb.apply_func(load_raw)
  X = data.drop(columns=[TARGET_COL]).skb.mark_as_X()
  y = data[TARGET_COL].skb.mark_as_y()

S5. Late mark_as_X when any feature is cross-row

The marker sits upstream of every cross-row step. Symptom: len(predictions) != n_predict_grid_rows, feature_steps=[], or a wrapper whose job is to filter NaNs the pipeline produced. Recovery: references/layer_examples.md. Do not loosen smoke.

A loader that computes y = col.shift(-H) and then marks X is S5 and S6. Horizon lives in Layer 2. Do not invent a wrapper estimator that shifts, dropnas, or filters nulls.

S6. Layer 1 doesn't know the question

Loaders describe what data exists. Horizon, lag, window, and task filters belong in Layer 2+. If an external consumer could not derive the step without knowing the task, push it past Layer 1.

Pre-flight — emit before any code

Pre-flight (build-ml-pipeline):
- [ ] sklearn, skrub, skore import
      Evidence: scratch/<ts>_check_tier1.py. Not inline python -c.
- [ ] api get for new skrub and sklearn symbols this turn
- [ ] Each skrub.var is a source id, not a baked path
- [ ] mark_as_X placement (loaded frame, or predict grid if cross-row)
- [ ] Layer 1 has no horizon, lag, or task filter
- [ ] cv on mark_as_X matches translation
      (date class | GroupKFold | KFold | no cv on holdout)
- [ ] data_dir_preview=None; no path literal in pipeline.py

Re-emit it with evidence before the final message.

References

  • references/snapshot.md — unfitted HTML, experiment cells, optional site build.
  • references/layer_examples.md — three layers. Read before proposing Layer 2.
  • references/source-binding.md — identifier versus materialized roots.
  • references/reproducibility_mechanics.md — Option 1 / 2 / 3.
  • references/common_patterns.md — tabular shapes with code.
  • references/custom-splitter.md — the split copies the locked deployment. The time series section is the date splitter when translation.scheme is date_time.
  • evaluate-ml-pipeline/references/metadata-routing.md — evaluate omits splitter=.
  • evaluate-ml-pipeline/references/custom-metrics.md — a score that is not a skore default.

Signals

GitHub stars
132
Forks
9
Last commit
Sep 2026

ahel review

  • K1binfo
    installs-packages (in evals/evals.json)
  • K1binfo
    installs-packages (in references/snapshot.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
build-ml-pipeline
Source
github.com/probabl-ai/skills