Monitor evaluation history

SkillMonitoring & ops

Review evaluation history, distinguish system, data, suite, and evaluator changes, and apply an operating response. Use for regressions, eval monitoring, release gates, or trend questions.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Monitor evaluation history skill

What this skill tells your AI

The instructions your AI receives, as published by ai-analyst-lab/ai-analyst in .claude/skills/monitor-evals/SKILL.md and read by ahel’s review.

Read versioned run manifests with helpers.evals.monitoring.load_history. Use classify_changes before comparing scores.

Separate:

  • system behavior changes;
  • data changes;
  • task-mix or suite-version changes;
  • evaluator changes; and
  • operational failures such as blocked tools or expired connections.

Compare only compatible runs. Show per-case and slice movement, not only the aggregate. Run frozen sentinel examples for model graders so evaluator drift does not look like system drift.

Apply the named operating rule:

  • continue when the intended change improved the target slice without a blocking regression;
  • investigate when the cause is unclear or several inputs changed;
  • rollback when a blocking regression follows a controlled system change; or
  • escalate when the evaluator, data, or authorization boundary may be invalid.

Record the owner and next action. Monitoring is an operating practice, not a dashboard someone passively observes.

Signals

GitHub stars
297
Forks
137
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
monitor-evals
Source
github.com/ai-analyst-lab/ai-analyst