Observability and reliability
SkillMonitoring & opsMakes systems debuggable and reliably operable, instrumentation, alerting that is worth waking for, service objectives, and learning from failure. Use this to instrument a service, fix alerting that is ignored, set error budgets or reliability targets, prepare for on-call, or run a blameless post-incident review.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Observability and reliability skill
What this skill tells your AI
The instructions your AI receives, as published by cbrock84/headcount in plugins/technology/skills/observability-and-reliability/SKILL.md and read by ahel’s review.
Monitoring tells you a thing you predicted is happening. Observability lets you ask a question you did not anticipate. Production failures are mostly the unanticipated kind.
Instrument for questions you have not thought of yet
Emit structured events with enough context to slice afterwards — request identifiers, user or tenant, version, dependency, outcome, duration. Free-text logs are unsearchable at volume and become expensive noise.
Propagate a correlation identifier across every hop. Without it, a distributed system is a set of independent stories and reconstructing one request is manual archaeology.
Measure what the user experiences at the percentile they experience it. A p50 latency graph is mostly a graph of the people who were not affected.
Alert on symptoms, not causes
Alert when users are affected or imminently will be. High CPU is not an alert; requests failing or slowing is. Cause-based alerting produces pages for conditions the system handled and no page for novel failures that hurt.
Every alert must be actionable, urgent and specific. If the recipient's honest response is to look and close it, delete the alert — it is training the on-call to ignore the page, and the ignored page is eventually the real one.
Alert fatigue is the actual reliability risk in most organizations. Fewer, better alerts beat coverage.
Objectives and error budgets
Set service level objectives from what users need, then treat the remainder as a budget to spend. This converts a sterile argument between shipping and stability into arithmetic: budget remaining means ship, budget exhausted means the next work is reliability.
Keep the internal objective tighter than any external commitment made through
operations:service-level-management, so you find out before the customer does.
Learn from incidents
Post-incident review exists to find what made the failure possible and hard to detect, not who touched it last. Human error is a starting question, never the finding: what made the error easy, and why did nothing catch it?
Track the time to detect separately from time to resolve. Long detection is an observability defect, and it is the part that repeats.
Produce a small number of real actions with owners and dates. A review generating fifteen actions generates none.
Sources
references/sources.md in this skill lists the outside authorities that settle the questions
here — what each one is authoritative for, and what you may do with it. Check them before
answering on anything they cover, and cite what you used. Most are free to read and not free
to reproduce; the use note on each is binding.
Tooling
Metrics and traces: Datadog, Grafana with Prometheus, New Relic, Honeycomb, and similar. Errors: Sentry, Rollbar, and similar. Logs: Elastic, OpenSearch, Loki, Splunk, and similar.
On-call and incident management: PagerDuty, Opsgenie, incident.io, FireHydrant, and similar.
Instrument with OpenTelemetry wherever you can. Vendor-specific instrumentation is the part that makes leaving expensive.
Never
- Page a human for something they cannot act on.
- Alert on a cause when you can alert on the symptom.
- Report reliability as an average when users experience the tail.
- Close an incident review with the finding that someone was careless.
Signals
- GitHub stars
- 2k
- Forks
- 262
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
observability-and-reliability- Source
- github.com/cbrock84/headcount
More in Monitoring & ops
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opsdashboard-builder
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & opseng-runbook
Skill · nexu-io
More in Monitoring & opsweekly-update
Skill · nexu-io
More in Monitoring & ops