SOC Metrics

SkillMonitoring & ops

Calculate and interpret SOC performance metrics, MTTD, MTTA, MTTR, MTTC, alert-to-incident ratio, false-positive and automation rates, and agent cost/latency, with maturity context; use for KPI reporting and SOC performance questions.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the SOC Metrics skill

What this skill tells your AI

The instructions your AI receives, as published by gensecaihq/wazuh-autopilot in backend/app/skills/soc-metrics/SKILL.md and read by ahel’s review.

Metrics should drive improvement, not vanity. Define precisely, compute consistently, and always show the trend and the sample size.

Time metrics (per incident, then median and p90 over the period)

MetricDefinitionStart → End
MTTD — mean time to detecthow long the attacker was active before detectionfirst malicious activity → first alert
MTTA — mean time to acknowledgetime until someone (agent or human) started workcase created → first triage finding
MTTC — mean time to containtime to stop the harmcase created → containment verified
MTTR — mean time to respond/resolvetime to close outcase created → resolved/closed

Report median and p90 alongside the mean — a few long incidents skew means. MTTD requires an incident timeline; mark it "n/a" where the start of malicious activity is unknown.

Volume and quality

MetricFormulaHealthy signal
Alert-to-incident ratioalerts ingested ÷ cases openedfalling over time with stable coverage
False-positive ratecases closed false_positive ÷ cases closedtrending down; > 50% means tuning needed
Escalation ratecases reaching investigation ÷ cases openedstable, explainable
Automation rateactions auto-approved or auto-executed ÷ all executed actionsincreases only with low rollback rate
Rollback rateactions rolled back ÷ executed< 5%; spikes mean over-aggressive autonomy
Verification rateexecuted actions verified ÷ executed→ 100%
Approval latencyaction proposed → approved (median)within policy expiry
Agent coverageactive Wazuh agents ÷ expected assets→ 100%
Telemetry lossalerts for rule 203 (event queue full) or rule 204 (queue flooded), and agents disconnected for > 1h (rule 504)zero; any non-zero means detection blind spots. Get detail from platform-engineer
MTTDfirst attacker event time (from the log, not the alert timestamp) → case openedfalling; ingestion delay shows up as the gap between event time and alert time

Agent (swarm) performance

Per agent: runs, error rate, median/p95 latency, tokens in/out, cost per run and per incident. Flag agents whose cost per incident rises without better outcomes, or whose error rate > 5%.

Interpretation

  • Always compare with the previous equal period and state the sample size.
  • A metric moving the "good" way can be bad: a falling FP rate plus falling incidents may mean a detection broke. Cross-check with alert volume and agent coverage.
  • SOC-CMM style maturity hints: consistent metric definitions, trend review in a regular meeting, and metrics tied to improvement actions indicate a managed (level 3+) process; ad-hoc or unreviewed metrics indicate initial/defined levels.

Output format

| Metric | This period | Previous | Change | n |

Follow with 3 bullet insights and 1–3 recommendations.

Output

  • save_report(kind="weekly"|"monthly", title, body_md) or include the table in the report being written by executive-reporting. standard_refs: NIST-CSF-2:ID.IM (improvement).

Signals

GitHub stars
57
Forks
16
Last commit
Sep 2026
Advanced
Item type
skill
Key
soc-metrics
Source
github.com/gensecaihq/wazuh-autopilot