Incident triage
SkillMonitoring & opsThe SRE team's runbook for triaging production latency and error-rate incidents. Use this whenever investigating an incident, a latency spike, elevated error rates, or when asked "what caused X" about a production service.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Incident triage skill
What this skill tells your AI
The instructions your AI receives, as published by thevibeworks/claude-code-docs in content/github/cwc-workshops/ship-your-first-managed-agent/incident-triage-runbook/SKILL.md and read by ahel’s review.
If you change the order below, say why in #sre.
Order of operations
- Pull deploys for the last 6h. Don't open the log first.
- Line the deploy timestamps up against
p99_latency_ms/error_ratefor the paged service. State the gap ("deploy 14:31, p99 moves 14:33"). - If a deploy lines up: pull the diff, read it. Check for the stuff in the next section.
- Then grep the log to confirm. Don't grep to fish.
- No deploy lines up → check
db_pool_utilizationacross checkout/cart/auth/inventory, then upstream deps.
Things that have burned us
In rough order of how often:
- per-row query where there used to be a batch
- cache decorator removed "temporarily"
- new query, no index
- blocking call in an async handler
- retry loop with no backoff
Write-up
One line at the bottom:
Root cause:
<sha>— one sentence on the mechanism.
If it wasn't a deploy, put the component or upstream dep where the sha goes (db-primary, stripe-api, whatever). Still one sentence.
Everything above that line is evidence. Keep it short; the long version goes in the postmortem doc.
Signals
- GitHub stars
- 42
- Forks
- 9
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
incident-triage-runbook- Source
- github.com/thevibeworks/claude-code-docs