Model-Assessment Skill
SkillMonitoring & opsUse when validating or evaluating a trained medical-imaging model. Audits split leakage and validation design, computes task-correct held-out metrics (Dice + HD95, AUROC + AUPRC, FROC, calibration), and covers uncertainty/OOD and Grad-CAM explainability, each with a gate.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Model-Assessment Skill skill
What this skill tells your AI
The instructions your AI receives, as published by aperivue/medsci-skills in skills/model-assessment/SKILL.md and read by ahel’s review.
Assess a trained medical-imaging model (in-house, vendor or open-weights; segmentation, classification or detection) in the order the evidence is built: Part A designs and audits the validation study, Part B computes task-correct held-out metrics, Part C adds the uncertainty / OOD / abstention layer a deployment claim needs, and Part D makes an explainability analysis survive review. Run the parts the request needs (a Grad-CAM question starts at Part D), but no metric headline is reported before the Part A split gate is green. Each part ends in a stdlib gate whose verdict is reproduced from a file — never report a pass without running it.
Numbers come only from code executed on the supplied predictions, split table, or the researcher's executed UQ/XAI code; if predictions or ground truth are missing, say so and stop. Integrate MONAI / nnU-Net, MAPIE, captum, pytorch-grad-cam and pretrained OOD scorers by reference — never reimplement them, never build or train the model, never run a model on real patient data.
Elsewhere: building/training → /model-scaffold; choosing or vetting the model →
/model-selection; data-stage preprocessing leakage → /imaging-data; paired model comparison /
added value over a baseline / decision curves / MRMC / ICC / calibration tables → /analyze-stats
(added value: incremental_value.md); AI-vs-expert reader rubric
and IRR → /design-ai-benchmarking; LLM/MLLM → /mllm-eval; general validity → /design-study;
classical-ML tabular calibration → /radiomics-ml; item-by-item reporting audit →
/check-reporting; a finished manuscript → /self-review or /peer-review (the MD0–MD11
model_development.md probe).
Part A — Validation design
The rationale behind Phases 2–7 — the full leakage taxonomy, the internal-vs-external ladder,
comparator design, run variance, test-set sizing, the reporting map — is in
${CLAUDE_SKILL_DIR}/references/validation_design.md (load on demand). Verify citations via
/search-lit (confirmed DOI/PMID), else mark [UNVERIFIED - NEEDS MANUAL CHECK]; flag an uncertain
CLAIM 2024 / TRIPOD+AI / Metrics Reloaded item [VERIFY] and ask.
Phase 1 — Task, intended-use horizon, analysis unit
State the task, the intended-use horizon (screening, triage, pre-procedure, post-hoc), the single headline metric the conclusion leans on, and the analysis unit (per-patient / per-lesion / per-image). Everything downstream is read against this; a per-lesion metric is never reported as per-patient.
Phase 2 — Leakage audit (run first)
Produce the emitted split-assignment table (patient_id,split) and run:
python3 ${CLAUDE_SKILL_DIR}/scripts/check_split_leakage.py \
--splits <split_assignment.csv> --out qc/split_leakage.json --strict
PATIENT_OVERLAP (a patient in ≥ 2 partitions) and MISSING_SEED (an unreproducible split) are
Major, SINGLE_PARTITION Minor — proven by set arithmetic on the ID column the gate prints as
id_col (pass --id-col with the patient identifier when it is not the one; an auto-picked
column whose name is not patient-level, such as image_id, gets a Minor ID_COL_NOT_PATIENT_LEVEL). A design with patient overlap is never
approved. Then walk the leakage the table cannot show (Kapoor & Narayanan, Patterns 2023):
preprocessing fit before the split (normalisation, resampling, foundation-model embeddings, ComBat
harmonisation over the whole cohort — /imaging-data gates the declared pipeline), site / scanner /
burned-in-label shortcuts, and temporal leakage (a random split where future and past coexist).
The decisive question: could any value used in training have been computed only with knowledge of
a test case?
Phase 3 — Validation tier
Classify honestly: apparent → internal random split → cross-validation → temporal → geographic / external (different site, scanner, vendor) → multi-site external. Cross-validation and bootstrap are development-time optimism corrections, not external validation. Flag a generalisability or deployment claim that outruns the design, and "external validation" where the single external set was used for tuning. Confirm the test set was touched once — no architecture search, hyperparameter sweep, early stopping, or threshold choice read it.
Phase 4 — Comparator
Clinical-only baseline, incremental value over an existing score, or reader comparison (hand the
rubric / inter-rater design to /design-ai-benchmarking).
Phase 5 — Test-set sizing
Count events per class in the test set, not the cohort total: a sparse positive set gives a CI
spanning much of the usable range, and calibration needs roughly ≥ 100 events. Hand formal sizing to
/calc-sample-size.
Phase 6 — Prospective evaluation and deployment-monitoring horizon
Retrospective external validation shows accuracy transfers, not that the model is safe and useful
in the workflow. For a clinical-use claim design the higher tier explicitly — silent / shadow
deployment → prospective comparative or impact study / RCT on a clinical endpoint →
post-deployment monitoring with recalibration-or-withdrawal triggers and subgroup audit
(references/validation_design.md §2b). Scope the claim to the tier reached: a retrospective
external study never claims deployment readiness or outcome benefit.
Phase 7 — Reporting-guideline fit
Map via /check-reporting: CLAIM 2024 (diagnostic imaging AI), TRIPOD+AI (prediction model),
STARD-AI (diagnostic accuracy), PROBAST+AI (risk of bias), and for a prospective/live
evaluation DECIDE-AI or CONSORT-AI / SPIRIT-AI.
Part B — Held-out metrics
Phase 8 — Compute task-correct metrics
Generate and execute evaluation code on the held-out predictions (Metrics Reloaded — Maier-Hein & Reinke et al., Nat Methods 2024):
- segmentation — Dice/IoU and a boundary metric (HD95 / NSD), per structure, 95% CIs by patient-level bootstrap (resample patients, not pixels or slices);
- classification — AUROC and AUPRC with patient-level bootstrap CIs, sensitivity/specificity, and PPV/NPV at the deployment prevalence, never bare accuracy on a balanced set. Report AUPRC with the test-set prevalence, which is its no-skill value: AUPRC moves with prevalence, so a value from an enriched or case-control test set does not carry over to deployment or across datasets. For multiclass, state the aggregation (one-vs-rest / macro / micro / pairwise / Obuchowski);
- detection — FROC / mAP with the IoU match criterion stated. Lesions and false positives cluster within patients, so CIs come from a patient-level bootstrap (resample patients, carrying all their lesions and false positives), not a Wilson/binomial interval over lesions, which is too narrow; compare two detectors' FROC curves with JAFROC (RJafroc), not per-lesion tests;
- interactive / promptable segmentation (SAM2 / MedSAM2 / nnInteractive) — the segmentation
metrics plus Dice-vs-interactions / number-of-clicks (NoC) to a target threshold,
initial-vs-converged (or peak) Dice, and per-case interaction/inference time. With two arms
(simulated prompting + human operator), record protocol fidelity — identical prompt types,
stopping rule, target threshold, seeds — because arm-to-arm comparability is what lets the human arm
validate the simulated one (human-arm design:
/design-study); - generative / synthesis — full-reference (MSE/RMSE/PSNR/SSIM) or no-reference (SNR/CNR, visual scores) quality plus a downstream-task evaluation: image quality is not clinical utility (Park et al., Radiol Med 2024);
- time-to-event discrimination (Harrell's C, time-dependent ROC) →
/analyze-stats.
Report the headline as the point estimate with a patient-level bootstrap 95% CI over the test
cases: that is the uncertainty of the test-set estimate. Seed-to-seed SD across training runs is a
different quantity (training-run variability, usually smaller) — report it separately, over ≥ 5
runs, for a training-recipe or model-comparison claim, and never present it as the CI. A frozen
vendor or open-weights model has no training runs to vary; its uncertainty is the test-set CI. Add
calibration — for a binary risk or diagnostic output, calibration-in-the-large (intercept), the
calibration slope and a flexible (loess) calibration curve, plus the Brier score; ECE only as a
supplementary top-label summary for multi-class confidence, with its binning stated — and
subgroup slices (the Model Card Factors). Emit results.md (metrics report) and a per-case
CSV for /analyze-stats. Load
${CLAUDE_SKILL_DIR}/references/metric_guide.md for the per-task checklist and
${CLAUDE_SKILL_DIR}/references/metric_selection_grounding.md for why each pairing is required and
the CLAIM 2024 fit map.
Phase 9 — Gate the metric choice
Declare the reported metrics in metrics_manifest.json (copy
${CLAUDE_SKILL_DIR}/templates/metrics_manifest.json; fields and allowed values in
references/metrics_manifest_schema.md), then:
python3 ${CLAUDE_SKILL_DIR}/scripts/check_metric_reporting.py \
--manifest metrics_manifest.json --out qc/metric_reporting.json --strict
PIXEL_ACCURACY_SEG / NO_BOUNDARY_METRIC / ACCURACY_ONLY / DETECTION_METRIC_MISSING /
INTERACTIVE_NO_INTERACTION_COUNT / GENERATIVE_NO_DOWNSTREAM / CLASSIFICATION_METRIC_MISSING /
SEGMENTATION_METRIC_MISSING (every Major) must be zero; the last two (no headline metric declared)
are manifest-only.
An off-list value exits 2; use "none" or "other:<description>". --report results.md --task <task> still runs the older keyword check on prose.
Known limits: manifest mode checks what is declared, not the reported numbers. Prose mode
(--report) tests keyword presence with a short negation window: "MSD" counts as mean surface
distance even when it names the Medical Segmentation Decathlon, "we did not compute the Hausdorff
distance or HD95" still counts HD95, "sensitivity and specificity were not reported" or "FROC was
not performed" still count as reported, a bare "map" ("saliency map") counts as mAP, and a wrapped
"mean average\nprecision" is not seen.
Part C — Uncertainty, OOD and selective prediction (deployment claims)
A deployment-framed model must say what it does when unsure or off-distribution. Read
${CLAUDE_SKILL_DIR}/references/uncertainty_guide.md for method choice and the manifest schema.
Phase 10 — Choose the uncertainty method, OOD guard and abstention rule
- Conformal (MAPIE) — prediction sets/intervals at nominal coverage; the strongest default with a calibration set. Its coverage guarantee is finite-sample but needs exchangeability, which can fail on clinical data, so measure achieved coverage on a test split disjoint from the calibration split and report it with its binomial CI — never report it as guaranteed.
- Deep ensemble — K ≥ 2 independent members (distinct seeds/inits); shared seeds underestimate epistemic uncertainty.
- MC-dropout — dropout active at inference, T passes; off, every pass is identical and the estimate collapses to a point prediction.
- Bayesian / last-layer Laplace — a light option.
- OOD guard — energy score, feature Mahalanobis, ODIN or max-softmax, evaluated on a held-out OOD set (different scanner/site/pathology) with detection AUROC and the operating point.
- Selective prediction — abstain at a pre-specified coverage/risk target; report the risk–coverage curve. A post-hoc threshold inflates accuracy-at-coverage.
- Under shift — report calibration/coverage on shifted or external data, not in-distribution only (Ovadia 2019).
Phase 11 — Emit and gate the uncertainty manifest
Write uncertainty_manifest.json:
{
"task": "classification",
"deployment_claim": true,
"uncertainty_method": "conformal",
"coverage_target": 0.90,
"coverage_validated": true,
"ood_method": "mahalanobis",
"ood_heldout_set": "external-ood-cohort",
"selective_prediction": true,
"selective_target": 0.95,
"calibration_under_shift": true
}
python3 ${CLAUDE_SKILL_DIR}/scripts/check_uncertainty_reporting.py --manifest uncertainty_manifest.json \
--out qc/uncertainty_reporting.json --strict
Verdicts: POINT_PREDICTION_NO_UNCERTAINTY, CONFORMAL_NO_COVERAGE_VALIDATION, OOD_NO_HELDOUT_SET
(Major); ENSEMBLE_NOT_INDEPENDENT, MCDROPOUT_DISABLED_AT_INFERENCE, SELECTIVE_NO_TARGET,
NO_CALIBRATION_UNDER_SHIFT (Minor). It audits the declared spec; it complements, not replaces,
Phase 8's executed calibration. Report TRIPOD+AI / DECIDE-AI deployment-monitoring fit via /check-reporting.
Part D — Explainability
A saliency / Grad-CAM map is the most over-interpreted artifact in imaging AI: Adebayo et al.
(NeurIPS 2018) showed many methods produce convincing maps independent of the model's weights and
labels. Read ${CLAUDE_SKILL_DIR}/references/explainability_guide.md for method by architecture,
sanity checks, localisation metrics and framing.
Phase 12 — Produce, sanity-check and quantify the maps
Choose the method for the architecture — Grad-CAM / Grad-CAM++ for CNNs, attention rollout for ViTs, integrated gradients / SHAP for attribution — wired through captum or pytorch-grad-cam. Run the Adebayo model-parameter and data (label) randomisation tests; a faithful map degrades when they are randomised, and both axes are the minimum bar. If the map is claimed to localise the finding, compute IoU / pointing game / Dice against ground-truth masks over the cohort — not eyeballed, cherry-picked cases. Frame a map as attribution ("where signal is attributed"), never as proof the model is correct or of causation.
Phase 13 — Emit and gate the explainability report
Write explainability_report.json:
{
"method": "grad-cam++",
"n_examples": 200,
"cohort_level": true,
"localization_metric": "iou",
"localization_value": 0.63,
"sanity_checks": ["model_randomization", "data_randomization"],
"interpretation": "localization"
}
interpretation: attribution / localization / faithfulness — never validation / causal.
python3 ${CLAUDE_SKILL_DIR}/scripts/check_explainability_report.py --manifest explainability_report.json \
--out qc/explainability_report.json --strict
Verdicts: SALIENCY_AS_VALIDATION, NO_SANITY_CHECK, NO_LOCALIZATION_METRIC (Major);
INSUFFICIENT_SANITY, CHERRY_PICKED_EXAMPLES, MISSING_METHOD (Minor).
Outputs and hand-off
- Part A: validation-design decision notes (leakage, tier, comparator, metric, sizing handoff, reporting
fit) and
qc/split_leakage.json. - Part B:
results.md, the per-case CSV, andqc/metric_reporting.json. - Part C:
uncertainty_manifest.json+qc/uncertainty_reporting.json. - Part D:
explainability_report.json+qc/explainability_report.json.
The per-case table → /analyze-stats (paired ΔAUC of frozen models on the same test patients —
DeLong or bootstrap; added value over a baseline per incremental_value.md; decision curves;
publication tables);
figures → /make-figures; numbers and subgroup performance → /model-card; Methods/Results →
/write-paper; compliance → /check-reporting; sizing → /calc-sample-size; the reviewer-side audit
of the draft → /self-review, whose ai_overclaiming / image_synthesis probes also check saliency
claims. Gate regression (${CLAUDE_SKILL_DIR}/): scripts/check_split_leakage_challenge/verify.sh,
scripts/metric_reporting_challenge/verify.sh, scripts/check_uncertainty_reporting_challenge/verify.sh,
scripts/check_explainability_report_challenge/verify.sh, and tests/test_*.sh.
Signals
- GitHub stars
- 320
- Forks
- 77
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
model-assessment- Source
- github.com/aperivue/medsci-skills
github.com/aperivue/medsci-skills
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonfastapi
Skill · fastapi
The pick for FastAPIintegration-fastapi
Skill · posthog
The pick for FastAPIinternal-comms
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & ops