Model-Assessment Skill

SkillMonitoring & ops

Use when validating or evaluating a trained medical-imaging model. Audits split leakage and validation design, computes task-correct held-out metrics (Dice + HD95, AUROC + AUPRC, FROC, calibration), and covers uncertainty/OOD and Grad-CAM explainability, each with a gate.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Model-Assessment Skill skill

What this skill tells your AI

The instructions your AI receives, as published by aperivue/medsci-skills in skills/model-assessment/SKILL.md and read by ahel’s review.

Assess a trained medical-imaging model (in-house, vendor or open-weights; segmentation, classification or detection) in the order the evidence is built: Part A designs and audits the validation study, Part B computes task-correct held-out metrics, Part C adds the uncertainty / OOD / abstention layer a deployment claim needs, and Part D makes an explainability analysis survive review. Run the parts the request needs (a Grad-CAM question starts at Part D), but no metric headline is reported before the Part A split gate is green. Each part ends in a stdlib gate whose verdict is reproduced from a file — never report a pass without running it.

Numbers come only from code executed on the supplied predictions, split table, or the researcher's executed UQ/XAI code; if predictions or ground truth are missing, say so and stop. Integrate MONAI / nnU-Net, MAPIE, captum, pytorch-grad-cam and pretrained OOD scorers by reference — never reimplement them, never build or train the model, never run a model on real patient data.

Elsewhere: building/training → /model-scaffold; choosing or vetting the model → /model-selection; data-stage preprocessing leakage → /imaging-data; paired model comparison / added value over a baseline / decision curves / MRMC / ICC / calibration tables → /analyze-stats (added value: incremental_value.md); AI-vs-expert reader rubric and IRR → /design-ai-benchmarking; LLM/MLLM → /mllm-eval; general validity → /design-study; classical-ML tabular calibration → /radiomics-ml; item-by-item reporting audit → /check-reporting; a finished manuscript → /self-review or /peer-review (the MD0–MD11 model_development.md probe).

Part A — Validation design

The rationale behind Phases 2–7 — the full leakage taxonomy, the internal-vs-external ladder, comparator design, run variance, test-set sizing, the reporting map — is in ${CLAUDE_SKILL_DIR}/references/validation_design.md (load on demand). Verify citations via /search-lit (confirmed DOI/PMID), else mark [UNVERIFIED - NEEDS MANUAL CHECK]; flag an uncertain CLAIM 2024 / TRIPOD+AI / Metrics Reloaded item [VERIFY] and ask.

Phase 1 — Task, intended-use horizon, analysis unit

State the task, the intended-use horizon (screening, triage, pre-procedure, post-hoc), the single headline metric the conclusion leans on, and the analysis unit (per-patient / per-lesion / per-image). Everything downstream is read against this; a per-lesion metric is never reported as per-patient.

Phase 2 — Leakage audit (run first)

Produce the emitted split-assignment table (patient_id,split) and run:

python3 ${CLAUDE_SKILL_DIR}/scripts/check_split_leakage.py \
  --splits <split_assignment.csv> --out qc/split_leakage.json --strict

PATIENT_OVERLAP (a patient in ≥ 2 partitions) and MISSING_SEED (an unreproducible split) are Major, SINGLE_PARTITION Minor — proven by set arithmetic on the ID column the gate prints as id_col (pass --id-col with the patient identifier when it is not the one; an auto-picked column whose name is not patient-level, such as image_id, gets a Minor ID_COL_NOT_PATIENT_LEVEL). A design with patient overlap is never approved. Then walk the leakage the table cannot show (Kapoor & Narayanan, Patterns 2023): preprocessing fit before the split (normalisation, resampling, foundation-model embeddings, ComBat harmonisation over the whole cohort — /imaging-data gates the declared pipeline), site / scanner / burned-in-label shortcuts, and temporal leakage (a random split where future and past coexist). The decisive question: could any value used in training have been computed only with knowledge of a test case?

Phase 3 — Validation tier

Classify honestly: apparent → internal random split → cross-validation → temporal → geographic / external (different site, scanner, vendor) → multi-site external. Cross-validation and bootstrap are development-time optimism corrections, not external validation. Flag a generalisability or deployment claim that outruns the design, and "external validation" where the single external set was used for tuning. Confirm the test set was touched once — no architecture search, hyperparameter sweep, early stopping, or threshold choice read it.

Phase 4 — Comparator

Clinical-only baseline, incremental value over an existing score, or reader comparison (hand the rubric / inter-rater design to /design-ai-benchmarking).

Phase 5 — Test-set sizing

Count events per class in the test set, not the cohort total: a sparse positive set gives a CI spanning much of the usable range, and calibration needs roughly ≥ 100 events. Hand formal sizing to /calc-sample-size.

Phase 6 — Prospective evaluation and deployment-monitoring horizon

Retrospective external validation shows accuracy transfers, not that the model is safe and useful in the workflow. For a clinical-use claim design the higher tier explicitly — silent / shadow deployment → prospective comparative or impact study / RCT on a clinical endpoint → post-deployment monitoring with recalibration-or-withdrawal triggers and subgroup audit (references/validation_design.md §2b). Scope the claim to the tier reached: a retrospective external study never claims deployment readiness or outcome benefit.

Phase 7 — Reporting-guideline fit

Map via /check-reporting: CLAIM 2024 (diagnostic imaging AI), TRIPOD+AI (prediction model), STARD-AI (diagnostic accuracy), PROBAST+AI (risk of bias), and for a prospective/live evaluation DECIDE-AI or CONSORT-AI / SPIRIT-AI.

Part B — Held-out metrics

Phase 8 — Compute task-correct metrics

Generate and execute evaluation code on the held-out predictions (Metrics Reloaded — Maier-Hein & Reinke et al., Nat Methods 2024):

  • segmentation — Dice/IoU and a boundary metric (HD95 / NSD), per structure, 95% CIs by patient-level bootstrap (resample patients, not pixels or slices);
  • classification — AUROC and AUPRC with patient-level bootstrap CIs, sensitivity/specificity, and PPV/NPV at the deployment prevalence, never bare accuracy on a balanced set. Report AUPRC with the test-set prevalence, which is its no-skill value: AUPRC moves with prevalence, so a value from an enriched or case-control test set does not carry over to deployment or across datasets. For multiclass, state the aggregation (one-vs-rest / macro / micro / pairwise / Obuchowski);
  • detection — FROC / mAP with the IoU match criterion stated. Lesions and false positives cluster within patients, so CIs come from a patient-level bootstrap (resample patients, carrying all their lesions and false positives), not a Wilson/binomial interval over lesions, which is too narrow; compare two detectors' FROC curves with JAFROC (RJafroc), not per-lesion tests;
  • interactive / promptable segmentation (SAM2 / MedSAM2 / nnInteractive) — the segmentation metrics plus Dice-vs-interactions / number-of-clicks (NoC) to a target threshold, initial-vs-converged (or peak) Dice, and per-case interaction/inference time. With two arms (simulated prompting + human operator), record protocol fidelity — identical prompt types, stopping rule, target threshold, seeds — because arm-to-arm comparability is what lets the human arm validate the simulated one (human-arm design: /design-study);
  • generative / synthesis — full-reference (MSE/RMSE/PSNR/SSIM) or no-reference (SNR/CNR, visual scores) quality plus a downstream-task evaluation: image quality is not clinical utility (Park et al., Radiol Med 2024);
  • time-to-event discrimination (Harrell's C, time-dependent ROC) → /analyze-stats.

Report the headline as the point estimate with a patient-level bootstrap 95% CI over the test cases: that is the uncertainty of the test-set estimate. Seed-to-seed SD across training runs is a different quantity (training-run variability, usually smaller) — report it separately, over ≥ 5 runs, for a training-recipe or model-comparison claim, and never present it as the CI. A frozen vendor or open-weights model has no training runs to vary; its uncertainty is the test-set CI. Add calibration — for a binary risk or diagnostic output, calibration-in-the-large (intercept), the calibration slope and a flexible (loess) calibration curve, plus the Brier score; ECE only as a supplementary top-label summary for multi-class confidence, with its binning stated — and subgroup slices (the Model Card Factors). Emit results.md (metrics report) and a per-case CSV for /analyze-stats. Load ${CLAUDE_SKILL_DIR}/references/metric_guide.md for the per-task checklist and ${CLAUDE_SKILL_DIR}/references/metric_selection_grounding.md for why each pairing is required and the CLAIM 2024 fit map.

Phase 9 — Gate the metric choice

Declare the reported metrics in metrics_manifest.json (copy ${CLAUDE_SKILL_DIR}/templates/metrics_manifest.json; fields and allowed values in references/metrics_manifest_schema.md), then:

python3 ${CLAUDE_SKILL_DIR}/scripts/check_metric_reporting.py \
  --manifest metrics_manifest.json --out qc/metric_reporting.json --strict

PIXEL_ACCURACY_SEG / NO_BOUNDARY_METRIC / ACCURACY_ONLY / DETECTION_METRIC_MISSING / INTERACTIVE_NO_INTERACTION_COUNT / GENERATIVE_NO_DOWNSTREAM / CLASSIFICATION_METRIC_MISSING / SEGMENTATION_METRIC_MISSING (every Major) must be zero; the last two (no headline metric declared) are manifest-only. An off-list value exits 2; use "none" or "other:<description>". --report results.md --task <task> still runs the older keyword check on prose.

Known limits: manifest mode checks what is declared, not the reported numbers. Prose mode (--report) tests keyword presence with a short negation window: "MSD" counts as mean surface distance even when it names the Medical Segmentation Decathlon, "we did not compute the Hausdorff distance or HD95" still counts HD95, "sensitivity and specificity were not reported" or "FROC was not performed" still count as reported, a bare "map" ("saliency map") counts as mAP, and a wrapped "mean average\nprecision" is not seen.

Part C — Uncertainty, OOD and selective prediction (deployment claims)

A deployment-framed model must say what it does when unsure or off-distribution. Read ${CLAUDE_SKILL_DIR}/references/uncertainty_guide.md for method choice and the manifest schema.

Phase 10 — Choose the uncertainty method, OOD guard and abstention rule

  • Conformal (MAPIE) — prediction sets/intervals at nominal coverage; the strongest default with a calibration set. Its coverage guarantee is finite-sample but needs exchangeability, which can fail on clinical data, so measure achieved coverage on a test split disjoint from the calibration split and report it with its binomial CI — never report it as guaranteed.
  • Deep ensemble — K ≥ 2 independent members (distinct seeds/inits); shared seeds underestimate epistemic uncertainty.
  • MC-dropout — dropout active at inference, T passes; off, every pass is identical and the estimate collapses to a point prediction.
  • Bayesian / last-layer Laplace — a light option.
  • OOD guard — energy score, feature Mahalanobis, ODIN or max-softmax, evaluated on a held-out OOD set (different scanner/site/pathology) with detection AUROC and the operating point.
  • Selective prediction — abstain at a pre-specified coverage/risk target; report the risk–coverage curve. A post-hoc threshold inflates accuracy-at-coverage.
  • Under shift — report calibration/coverage on shifted or external data, not in-distribution only (Ovadia 2019).

Phase 11 — Emit and gate the uncertainty manifest

Write uncertainty_manifest.json:

{
  "task": "classification",
  "deployment_claim": true,
  "uncertainty_method": "conformal",
  "coverage_target": 0.90,
  "coverage_validated": true,
  "ood_method": "mahalanobis",
  "ood_heldout_set": "external-ood-cohort",
  "selective_prediction": true,
  "selective_target": 0.95,
  "calibration_under_shift": true
}
python3 ${CLAUDE_SKILL_DIR}/scripts/check_uncertainty_reporting.py --manifest uncertainty_manifest.json \
  --out qc/uncertainty_reporting.json --strict

Verdicts: POINT_PREDICTION_NO_UNCERTAINTY, CONFORMAL_NO_COVERAGE_VALIDATION, OOD_NO_HELDOUT_SET (Major); ENSEMBLE_NOT_INDEPENDENT, MCDROPOUT_DISABLED_AT_INFERENCE, SELECTIVE_NO_TARGET, NO_CALIBRATION_UNDER_SHIFT (Minor). It audits the declared spec; it complements, not replaces, Phase 8's executed calibration. Report TRIPOD+AI / DECIDE-AI deployment-monitoring fit via /check-reporting.

Part D — Explainability

A saliency / Grad-CAM map is the most over-interpreted artifact in imaging AI: Adebayo et al. (NeurIPS 2018) showed many methods produce convincing maps independent of the model's weights and labels. Read ${CLAUDE_SKILL_DIR}/references/explainability_guide.md for method by architecture, sanity checks, localisation metrics and framing.

Phase 12 — Produce, sanity-check and quantify the maps

Choose the method for the architecture — Grad-CAM / Grad-CAM++ for CNNs, attention rollout for ViTs, integrated gradients / SHAP for attribution — wired through captum or pytorch-grad-cam. Run the Adebayo model-parameter and data (label) randomisation tests; a faithful map degrades when they are randomised, and both axes are the minimum bar. If the map is claimed to localise the finding, compute IoU / pointing game / Dice against ground-truth masks over the cohort — not eyeballed, cherry-picked cases. Frame a map as attribution ("where signal is attributed"), never as proof the model is correct or of causation.

Phase 13 — Emit and gate the explainability report

Write explainability_report.json:

{
  "method": "grad-cam++",
  "n_examples": 200,
  "cohort_level": true,
  "localization_metric": "iou",
  "localization_value": 0.63,
  "sanity_checks": ["model_randomization", "data_randomization"],
  "interpretation": "localization"
}

interpretation: attribution / localization / faithfulness — never validation / causal.

python3 ${CLAUDE_SKILL_DIR}/scripts/check_explainability_report.py --manifest explainability_report.json \
  --out qc/explainability_report.json --strict

Verdicts: SALIENCY_AS_VALIDATION, NO_SANITY_CHECK, NO_LOCALIZATION_METRIC (Major); INSUFFICIENT_SANITY, CHERRY_PICKED_EXAMPLES, MISSING_METHOD (Minor).

Outputs and hand-off

  • Part A: validation-design decision notes (leakage, tier, comparator, metric, sizing handoff, reporting fit) and qc/split_leakage.json.
  • Part B: results.md, the per-case CSV, and qc/metric_reporting.json.
  • Part C: uncertainty_manifest.json + qc/uncertainty_reporting.json.
  • Part D: explainability_report.json + qc/explainability_report.json.

The per-case table → /analyze-stats (paired ΔAUC of frozen models on the same test patients — DeLong or bootstrap; added value over a baseline per incremental_value.md; decision curves; publication tables); figures → /make-figures; numbers and subgroup performance → /model-card; Methods/Results → /write-paper; compliance → /check-reporting; sizing → /calc-sample-size; the reviewer-side audit of the draft → /self-review, whose ai_overclaiming / image_synthesis probes also check saliency claims. Gate regression (${CLAUDE_SKILL_DIR}/): scripts/check_split_leakage_challenge/verify.sh, scripts/metric_reporting_challenge/verify.sh, scripts/check_uncertainty_reporting_challenge/verify.sh, scripts/check_explainability_report_challenge/verify.sh, and tests/test_*.sh.

Signals

GitHub stars
320
Forks
77
Last commit
Sep 2026
Advanced
Item type
skill
Key
model-assessment
Source
github.com/aperivue/medsci-skills