Severity Scoring

SkillDev tools

Compute a defensible priority (P1, P4), severity and confidence for a case from rule level, asset criticality, blast radius and ATT&CK tactic; use whenever you set or change case severity.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Severity Scoring skill

What this skill tells your AI

The instructions your AI receives, as published by gensecaihq/wazuh-autopilot in backend/app/skills/severity-scoring/SKILL.md and read by ahel’s review.

Severity must be explainable. Use this formula, then sanity-check it.

Inputs (each scored 0–10)

FactorWeightHow to score
Rule level (L)0.30min(10, rule.level * 10 / 15) using the highest level in the case (see note below)
Asset criticality (A)0.2510 crown-jewel / domain controller / prod DB; 7 prod server; 4 internal workstation; 2 lab/test
Tactic weight (T)0.20see table below, highest tactic present
Blast radius (B)0.151 host = 2; 2–5 hosts = 5; >5 hosts or domain-wide identity = 9; org-wide = 10
Evidence strength (E)0.1010 confirmed success (e.g. login after brute force, file dropped); 5 attempt; 2 anomaly only

score = 0.30L + 0.25A + 0.20T + 0.15B + 0.10E

About Wazuh levels

Wazuh levels describe the rule, not your environment. Some are misleading in context:

  • High-level rules can be environmental or operational. For example, rule 204 (agent event queue flooded, level 12) is a pipeline problem, and rule 1003 (oversized syslog message, level 13) is often a misbehaving device. Keep L from the level, but score E low unless there is attacker evidence.
  • Low-level rules can be the key evidence. For example, rule 5715 (sshd authentication success, level 3) right after rule 5712 (brute force, level 10) from the same source is a likely compromise. Score E = 10 and use the highest level in the case, not the success event's.
  • Vulnerability-detector levels follow CVE severity (rules 23503, 23504, 23505, 23506 = Low, Medium, High, Critical). That is exposure, not activity: route to vulnerability-prioritization instead of scoring it as an intrusion.

Tactic weights (MITRE ATT&CK)

TacticWeight
Reconnaissance, Resource Development2
Initial Access (attempt), Discovery4
Execution, Persistence, Defense Evasion6
Credential Access, Privilege Escalation7
Lateral Movement, Command and Control8
Collection, Exfiltration9
Impact10

Mapping

ScorePrioritySeverityTarget response
≥ 8.0P1criticalimmediate, page on-call
6.0–7.9P2highwithin 1 hour
4.0–5.9P3mediumsame business day
< 4.0P4low / informationalbacklog / batch

Overrides (apply after the formula):

  • Confirmed data exfiltration or ransomware behaviour → P1 regardless.
  • Only reconnaissance with no success, external source → cap at P3.
  • Asset tagged protected/critical in org context → floor at P2.

Confidence (0–1)

Confidence is how sure you are the activity is malicious, separate from severity.

EvidenceConfidence
Multiple independent signals + threat-intel match0.9–1.0
Clear pattern (e.g. brute force then success)0.75–0.9
Single high-fidelity rule0.6–0.75
Single generic rule / anomaly0.3–0.6
Likely benign, unexplained< 0.3

perform_risk_assessment may be used as an extra input; don't let it override your own reasoning without explanation.

Output

  • update_case(case_id, severity=..., confidence=...).
  • add_finding titled Severity rationale containing the factor table with scores, the computed score, overrides applied, priority, and confidence reasoning. standard_refs: NIST-800-61r3.

Signals

GitHub stars
57
Forks
16
Last commit
Sep 2026
Advanced
Item type
skill
Key
severity-scoring
Source
github.com/gensecaihq/wazuh-autopilot