Incident Postmortem

SkillMonitoring & ops

Write blameless incident postmortems following Google SRE methodology. Covers timeline reconstruction, root cause analysis (5 Whys), impact assessment, and action items with owners. Use after production incidents, outages, or significant bugs to prevent recurrence.

Use Incident Postmortem in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Incident Postmortem and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Incident Postmortem skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Incident PostmortemStart free

What this skill tells your AI

The instructions your AI receives, as published by hoavdc/codexkit in skills/codexkit-incident-postmortem/SKILL.md and read by Ahel’s review.

When to Use

  • After any production incident that affected users
  • After near-misses that revealed systemic weaknesses
  • When establishing a blameless postmortem culture
  • When leadership requests incident RCA (Root Cause Analysis)

Procedure

Step 1 — Incident Summary

Write a 2–3 sentence summary:

  • What happened, when, and how long
  • Impact in business terms (users affected, revenue lost, SLA breach)
  • Current status (resolved, monitoring, partially mitigated)

Severity classification:

SeverityCriteria
SEV-1Complete outage or data loss affecting all users
SEV-2Major feature unavailable or significant degradation
SEV-3Minor feature impact, workaround available
SEV-4Cosmetic or non-user-facing issue

Step 2 — Timeline

Build a minute-by-minute timeline:

Time (UTC)EventSource
14:02Deploy v2.3.1 to productionCI/CD
14:05Error rate spikes to 15%Datadog
14:08PagerDuty alerts on-call engineerPagerDuty
14:12On-call begins investigationSlack thread
14:25Root cause identified: DB migration timeoutLogs
14:30Rollback initiatedCI/CD
14:35Service fully recoveredMonitoring

Time to Detect (TTD): 3 min | Time to Resolve (TTR): 33 min

Step 3 — Root Cause Analysis (5 Whys)

  1. Why did the service fail? → Database queries timed out
  2. Why did queries time out? → Migration locked a critical table for 8 minutes
  3. Why was the table locked? → ALTER TABLE ran without CONCURRENTLY flag
  4. Why wasn't it concurrent? → Migration script didn't follow the safe-migration checklist
  5. Why wasn't the checklist followed? → No automated check in CI pipeline

Root cause: Missing CI check for safe migration patterns.

Step 4 — Impact Assessment

MetricValue
Duration33 minutes
Users affected~12,000 (8% of DAU)
Requests failed~45,000 (500 errors)
Revenue impact~$2,400 estimated
SLA impact99.92% (target: 99.95%) — SLA breached

Step 5 — What Went Well / What Didn't

What Went WellWhat Didn't Go Well
Fast detection (3 min TTD)No pre-deploy migration testing
Clear rollback procedureMigration script wasn't reviewed
Team communicated via war roomPagerDuty escalation was slow

Step 6 — Action Items

#ActionTypeOwnerPriorityDeadline
1Add CI check for safe migration patternsPreventPlatformP1Sprint 5
2Add migration dry-run to staging deployDetectDevOpsP1Sprint 5
3Improve PagerDuty escalation timingProcessSREP2Sprint 6

Action types: Prevent (stop recurrence), Detect (catch faster), Mitigate (reduce impact)

Inputs

InputRequiredFormat
Incident descriptionYesWhat happened, when
Timeline / log dataYesTimestamps with events
Impact dataYesUsers affected, duration, revenue
Team notesRecommendedSlack threads, war room notes

Output

## Postmortem — [Incident Title]

**Date:** 2024-03-15 | **Severity:** SEV-2 | **Duration:** 33 min
**Author:** [On-call engineer] | **Reviewers:** [Team leads]

### Summary
Production database queries timed out for 33 minutes due to a table-locking
migration, affecting ~12,000 users and breaching our 99.95% SLA target.

### Timeline
[Minute-by-minute timeline]

### Root Cause
[5 Whys chain → root cause]

### Impact
[Quantified impact table]

### Lessons Learned
[What went well / what didn't]

### Action Items
[Prioritized actions with owners and deadlines]

Definition of Done

  • Blameless language throughout (no individual blame)
  • Timeline with timestamps and sources
  • 5 Whys reaching a systemic root cause
  • Impact quantified (users, duration, revenue, SLA)
  • Action items typed (prevent/detect/mitigate) with owners

Quality Criteria

  • Steps are executable in sequence without external context
  • Decision points have clear if/then branching
  • Rollback or abort procedures are documented for risky steps
  • Expected duration or time-per-step is estimated

Verification (4C)

CheckQuestion
CorrectnessDo the steps execute correctly in the order specified?
CompletenessAre decision points, error handling, and escalation paths all documented?
Context-fitCould someone with the right access but no prior context complete this runbook?
ConsequenceIf Step N fails and the operator skips to Step N+1, what breaks?

Edge Cases

  • Steps require access the operator doesn't have — Document exact access requirements upfront. Include escalation contact for emergency access.
  • Environment differs from documented state — Add a pre-flight check as Step 0 to verify prerequisites before starting.
  • Runbook is triggered during off-hours — Document who to contact and which steps can be safely deferred to business hours.

Changelog

  • v1.0.0 — Initial release

Signals

GitHub stars
25
Forks
13
Last commit
Oct 2026
Advanced
Item type
skill
Key
codexkit-incident-postmortem
Source
github.com/hoavdc/codexkit