Alert Root Cause Analysis

SkillMonitoring & ops

Performs root cause analysis on active alerts, correlating metrics and inspection records.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Alert Root Cause Analysis skill

What this skill tells your AI

The instructions your AI receives, as published by kubehan/promai in skills/alert-root-cause/SKILL.md and read by ahel’s review.

当用户询问告警原因、触发根因、或要求分析某个告警时,使用此工作流。

前置条件

用户通常会说 "分析一下这个告警" 或 "为什么 CPU 告警了"。先使用已知的 metric_name 或从上下文推断。

步骤

1. 使用 analyze_alert 工具

这是分析告警的核心工具。直接传入指标名称:

analyze_alert(metric_name="node_cpu_seconds_total")

如果用户提到了具体实例,带上 instance:

analyze_alert(metric_name="node_memory_MemAvailable_bytes", instance="192.168.1.100:9100")

2. 查询关联指标

analyze_alert 返回后,根据结果补充查询关联指标:

如果是 CPU 告警,检查负载:

node_load1

如果是内存告警,检查 OOM:

(node_memory_Active_bytes / node_memory_MemTotal_bytes) * 100

如果是磁盘告警,检查 inode:

(node_filesystem_files_free{fstype!=""} / node_filesystem_files{fstype!=""}) * 100

3. 查看巡检记录

使用 list_reports 获取近期巡检报告:

list_reports(status="problem", page_size=5)

如果找到相关报告,用 get_report_detail 查看详情:

get_report_detail(report_id=123)

4. 汇总分析

  • 对比告警触发前后的指标变化趋势
  • 结合巡检报告中的异常发现
  • 给出根因判断和处理建议
  • 如果确认是严重问题,使用 push_report 通知值班人员

常见场景示例

CPU 持续高负载

  1. analyze_alert(metric_name="node_cpu_seconds_total")
  2. query_metrics(promql="topk(5, sum by (process) (rate(node_procs_running{}[5m])))")
  3. 建议使用 exec 确认是否有异常进程

Signals

GitHub stars
124
Forks
44
Last commit
Sep 2026
Advanced
Item type
skill
Key
alert-root-cause
Source
github.com/kubehan/promai