Discovery Skill

SkillDev tools

Use this skill when running systematic quality discovery and issue detection. Runs modular probes adapted to the project's tech stack, presents findings interactively for user triage, and creates VCS issues for confirmed problems. Invoked standalone via /discovery or embedded in session-end.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Discovery Skill skill

What this skill tells your AI

The instructions your AI receives, as published by kanevry/session-orchestrator in skills/discovery/SKILL.md and read by ahel’s review.

Invocation Modes

Two modes of operation:

  • Standalone (/discovery [scope]): Full 6-phase flow with interactive triage (Phases 0-6)
  • Embedded (from session-end when discovery-on-close: true): Phases 0-4 only, returns structured findings to session-end

The scope argument accepts: all (default), code, infra, ui, arch, session, audit, vault, feature, or comma-separated like code,session.

Phase 0: Bootstrap Gate

Read skills/_shared/bootstrap-gate.md and execute the gate check. If the gate is CLOSED, invoke skills/bootstrap/SKILL.md and wait for completion before proceeding. If the gate is OPEN, continue to Phase 1.

Phase 1: Read Session Config

Read and parse Session Config per skills/_shared/config-reading.md. Store result as $CONFIG.

Discovery-relevant fields (parse these specifically):

  • discovery-on-close, discovery-probes, discovery-exclude-paths, discovery-severity-threshold, discovery-confidence-threshold, discovery-parallelism
  • test-command, typecheck-command, lint-command
  • pencil, vcs, cross-repos, stale-issue-days

Phase 2: Stack Detection & Probe Activation

Detect the project's tech stack via marker file checks. Use Glob and run checks in parallel:

Marker File(s)Activates
package.jsonJS/TS probes
tsconfig.jsonTypeScript probes
requirements.txt / pyproject.tomlPython probes
Dockerfile / docker-compose.ymlContainer probes
vercel.json / .vercel/Vercel probes
.github/workflows/GitHub CI probes
.gitlab-ci.ymlGitLab CI probes
supabase/Supabase probes
next.config.* / nuxt.config.*SSR probes
tailwind.config.*Tailwind probes
Pencil in Session Configdesign-drift probe
.orchestrator/bootstrap.lockharness-audit probe
.vault.yaml OR Session Config vault-integration.enabled: truevault probes
package.json / requirements.txt / Cargo.toml AND Session Config slopcheck.enabled: true AND slopcheck.sources includes "discovery"supply-chain probe (skills/discovery/probes-supply-chain.md)
docs/ directory present AND Session Config docs-staleness.enabled: truedocs-staleness probe (skills/discovery/probes-docs.md)
CLAUDE.md (or AGENTS.md on Codex CLI) or README.md present in repo rootssot-code-diff probe (skills/discovery/probes-docs.md) — always active within the docs category, no Session Config gate

Build Activation Set

  1. Start with all probes whose marker files are present
  2. If discovery-probes is set in config, intersect with that list
  3. If a scope argument was passed, restrict to that category
  4. Remove probes whose activation conditions are not met

The audit probe activates when bootstrap.lock is present OR when discovery-probes config explicitly lists audit.

The vault probe activates when .vault.yaml is present in the repo root OR when vault-integration.enabled: true in Session Config OR when discovery-probes config explicitly lists vault.

The feature probe activates ONLY when the scope argument includes feature OR when discovery-probes config explicitly lists feature. It never activates under bare all — feature discovery is an explicitly requested scan, and its interactive router (below) cannot run in embedded mode.

Exclude Paths

Default exclude paths (always apply):

  • node_modules/, .git/, dist/, build/, .next/, .nuxt/, coverage/

Add any paths from discovery-exclude-paths in Session Config.

VCS Detection

VCS Reference: Detect the VCS platform per the "VCS Auto-Detection" section of the gitlab-ops skill.

Status Report

Report: "Discovery: [N] probes active across [categories]. Stack: [detected]. Threshold: [severity]."

Feature-Scope Router (standalone only)

Fires ONLY when the active scope set includes feature AND the skill is running standalone (coordinator context). In embedded mode (session-end dispatches discovery as an Explore subagent — AskUserQuestion is unavailable in subagents per .claude/rules/ask-via-tool.md AUQ-004), the router MUST NOT fire; proceed directly with the grounded scan as the non-interactive default.

Judgment-based PM work (opportunity framing, personas, market-sizing) deliberately stays OUT of the verified-findings pipeline (Epic #750) — this router is the fork point that keeps grounded, evidence-anchored discovery separate from open-ended product judgment.

When the gate above is satisfied, present exactly this AskUserQuestion (AUQ-003 shape) before proceeding to Phase 3:

AskUserQuestion({
  questions: [{
    question: "Scope `feature` was requested. How should this run handle it?",
    header: "Scope",
    options: [
      { label: "Grounded scan (Recommended)", description: "Every finding is tied to a file and line and is verified before it can become an issue. Cost: two extra probes (intent-drift, stubbed-dead-feature)." },
      { label: "Also judgment topics", description: "Same scan, plus open product questions (opportunity framing, personas) kept as notes. They never become issues; after Phase 5 you pick where they go." },
      { label: "Route out", description: "No scan at all. You get a pointer to /brainstorm (product ideation) or /grill (assumption stress-test) instead." },
      { label: "Skip", description: "Drops `feature` (the probes for half-built and drifted features) from this run; the other scopes still run." }
    ],
    multiSelect: false
  }]
})

On user selection (every branch is explicit — do not fall through):

  • "Grounded scan" → keep feature in the active scope set and continue to Phase 3 unchanged.
  • "Also judgment topics" → keep feature in the active scope set and continue to Phase 3 unchanged for the grounded probes. Collect any judgment-based product questions (opportunity framing, personas, market-sizing) encountered during the scan as plain notes — do not route them through Phase 4 (Verification) or Phase 6 (Issue-Creation) at any point. AFTER Phase 5 triage completes, present a SECOND AskUserQuestion that forks how the collected topics are handled:
AskUserQuestion({
  questions: [{
    question: "Where should the collected judgment topics go?",
    header: "Topics",
    options: [
      { label: "Inline synthesis (Recommended)", description: "Sketches an outcome/persona pass into `### Judgment Topics (non-verified)` (a report section that never becomes issues). Cost: no second run." },
      { label: "Route to /brainstorm", description: "Hands the topics to /brainstorm as its opening context, for a full question-and-answer design dialogue." },
      { label: "Route to /plan feature", description: "Hands the topics to /plan feature as its opening context, for feature-PRD scoping." }
    ],
    multiSelect: false
  }]
})
  • "Inline synthesis" → append a ### Judgment Topics (non-verified) section to the discovery report. For each collected topic, write a brief candidate outcome/opportunity sketch (Teresa Torres OST framing) and, where a persona angle is evident, a one-line persona note — every line explicitly labeled (non-verified sketch). This is prose synthesis, never a probe finding.
  • "Route to /brainstorm" → append the same ### Judgment Topics (non-verified) section, but render each topic as a one-line pointer plus the recommendation to invoke /brainstorm with the collected topic context pre-filled as its initial prompt. Do not auto-invoke /brainstorm — surface the recommendation and let the operator trigger it.
  • "Route to /plan feature" → same rendering as the /brainstorm branch above, pointing at /plan feature with the collected topic context pre-filled instead.

Invariant (PRD §4 Non-Goals) — holds across all three sub-branches above: judgment items are report-appendix ONLY — they MUST NOT become findings, MUST NOT pass through Phase 4 verification, and MUST NOT produce issues in Phase 6.

  • "Route out" → remove feature from the active scope set. Do NOT run the feature probes. Recommend the operator invoke /brainstorm or /grill after this run (name whichever fits the operator's question), then continue to Phase 3 with the remaining scopes — or, if feature was the ONLY requested scope, stop after reporting: "Feature scope routed out — no probes to run. Invoke /brainstorm or /grill directly."
  • "Skip" → remove feature from the active scope set before Phase 3. If feature was the only requested scope, stop with: "Feature scope skipped — nothing to run."

Phase 3: Probe Execution

--since Filtering (when since_ref is provided)

When since_ref is set (passed from the /discovery --since <git-ref> invocation):

  1. Call changedFilesSince(since_ref) from scripts/lib/discovery/helpers.mjs.
  2. If the helper throws (ref unresolvable), surface the error to the user and halt.
  3. If the result is [] (no files changed since the ref), emit:
    No files changed since <since_ref>. Skipping discovery.
    
    and exit with status 0. Do NOT fall back to a full-repo scan.
  4. If the result is a non-empty array, pass it as changedFiles context to each probe agent below.

Probe exemptions: The vault-staleness probe and the harness-audit probe are EXEMPT from --since filtering — they always scan the full repository because their analysis targets metadata (vault narrative staleness, bootstrap lock state) that is not file-diff-gated. This exemption is advisory: no code enforcement is applied in this wave. The probe agents will naturally read whole-repo state; the changedFiles context they receive from --since is informational and does not restrict their glob/grep scope.

Dispatch probe agents IN PARALLEL using the Agent tool. Group by category (max $CONFIG['discovery-parallelism'] agents, default 5):

Cursor IDE: No Agent() tool available. Run probes sequentially within the current session — one category at a time. Complete each category's analysis before moving to the next.

  • Code probes agent: Runs all activated code probes (hardcoded-values, orphaned-annotations, dead-code, ai-slop, type-safety-gaps, test-coverage-gaps, test-anti-patterns, security-basics)
  • Infra probes agent: Runs all activated infra probes
  • UI probes agent: Runs all activated UI probes
  • Arch probes agent: Runs all activated arch probes
  • Session probes agent: Runs all activated session probes
  • Audit probes agent: Runs harness-audit probe
  • Vault probes (skills/discovery/probes-vault.md): invokes skills/discovery/probes/vault-staleness.mjs and skills/discovery/probes/vault-narrative-staleness.mjs directly via node. Each probe returns {findings, metrics, duration_ms}. The runner reports FINDING: blocks per finding and appends summary records to .orchestrator/metrics/vault-staleness.jsonl and vault-narrative-staleness.jsonl.
  • Supply-chain probe (skills/discovery/probes-supply-chain.md): invokes skills/discovery/probes/supply-chain-slopcheck.mjs directly via node. Gated: only activates when slopcheck.enabled: true AND "discovery" is in slopcheck.sources (Session Config). The probe returns {findings, summary}. SLOP findings surface as critical, ASSUMED as medium, LEGITIMATE packages generate no finding. See probes-supply-chain.md for invocation details and classification reference.
  • Docs probe (skills/discovery/probes-docs.md): invokes skills/discovery/probes/docs-staleness.mjs AND skills/discovery/probes/ssot-code-diff.mjs directly via node. docs-staleness.mjs is gated: only activates when a docs/ directory exists AND docs-staleness.enabled: true (Session Config); it scans docs/*.md (root level) and docs/examples/*.md for filesystem-mtime staleness against the docs-staleness.thresholds.living threshold (default 90d) — docs/adr/ and docs/prd/ are deliberately excluded. ssot-code-diff.mjs is ungated (no Session Config key) — it always runs when CLAUDE.md or README.md is present, diffing hardcoded doc "count" claims (e.g. "(13 rules)") against the live code/filesystem value they describe. Both probes return {findings, metrics, duration_ms}. See probes-docs.md for invocation details and severity escalation.
  • Feature probes agent (skills/discovery/probes-feature.md): intent-drift + stubbed-dead-feature + feature-request-cluster probes. The first two MUST carry file_path:line_number anchors that survive the Phase 4.2 re-read (±3 lines); feature-request-cluster instead anchors on a VCS issue-ID set and MUST carry verification_method: vcs-issue — see Phase 4.2's dual verification path below.

Each agent receives:

  • The probe definitions from probes-intro.md (confidence scoring reference) AND the category-specific probes-<category>.md file for this agent's category (include the actual grep commands/patterns in the prompt)
  • The exclude paths list
  • The project root path
  • When since_ref was provided and changedFiles is non-empty: the changedFiles array (informational context for per-probe filtering — per-probe filtering enforcement is deferred to W3)
  • Tools: Read, Grep, Glob, Bash (read-only -- no Edit/Write)
  • Instruction: "Run each probe. For each finding, output EXACTLY this format:"
FINDING:
  probe: <probe_name>
  category: <category>
  severity: <critical|high|medium|low>
  file_path: <absolute path>
  line_number: <number>
  matched_text: <exact text from tool output>
  title: <short title for the finding>
  description: <1-2 sentence description>
  recommended_fix: <concrete fix suggestion>

"If a probe's activation condition is not met, skip it with: SKIPPED: <probe_name> -- " "If a probe command fails, skip it with: FAILED: <probe_name> -- " "Do NOT fabricate findings. Only report what tool output confirms."

CRITICAL: run_in_background: false for all agents.

Skip categories with no activated probes (don't dispatch empty agents).

Phase 4: Verification & Scoring

After all probe agents complete:

4.1 Parse Findings

Collect all FINDING: blocks from agent outputs into a unified findings list.

4.2 Verification Pass

For EACH finding, use the verification path matching how the finding is anchored:

Default path — file-anchored findings (no verification_method field, or verification_method: file-line):

  1. Read the file at file_path:line_number using the Read tool
  2. Confirm matched_text appears at or near that line (+/-3 lines tolerance)
  3. If NOT confirmed, discard with note "false positive -- text not found at reported location"

VCS-issue-anchored findings (verification_method: vcs-issue) — e.g. feature-request-cluster, whose findings anchor on a set of issue IDs rather than a file location:

  1. For each issue ID in the finding's Location set, verify it via the VCS CLI — syntax reference: skills/gitlab-ops/SKILL.md § "Common CLI Commands" (glab issue view <IID> / gh issue view <NUMBER>); do not duplicate the CLI syntax here

  2. Confirm the issue exists and is still open (state != closed)

  3. If an issue ID no longer resolves or was closed since detection, drop it from the Location set; if EVERY issue in the set is gone/closed, discard the finding entirely with note "false positive -- all referenced issues closed/removed since detection"

  4. Report: "Verification: N confirmed, M discarded as false positives"

4.2a Confidence Scoring

For each verified finding, assign a confidence score (0-100) based on three factors:

FactorLow (+0)Medium (+10)High (+20)
Pattern specificityGeneric match (URL, TODO)Moderate (orphaned annotation, magic number)Specific (API key regex, eval(), SQL injection)
File contextTest fixture, example, seed data, docsUtility, config, scriptsProduction source, API handler, middleware
Historical signalPreviously dismissed as false positiveNo prior data (first occurrence)Recurring issue (confirmed in learnings.jsonl)

Scoring rules:

  1. Start at 40 (baseline)
  2. Add each factor's score (+0/+10/+20)
  3. Clamp to 0-100 range
  4. Critical severity override: findings with severity critical get a minimum confidence of 70 — they are NEVER auto-deferred

Threshold: Read discovery-confidence-threshold from Session Config (default: 60). If not configured, use 60.

Annotate each finding with its confidence score for Phase 5 presentation.

4.3 Deduplication

Two findings are duplicates if:

  • Same file_path AND
  • Overlapping line range (+/-5 lines) AND
  • Different probes

Keep the higher severity finding. Merge descriptions.

4.4 Apply Thresholds

  1. Severity filter: Remove findings below discovery-severity-threshold from Session Config.
  2. Confidence filter: Remove findings with confidence score below discovery-confidence-threshold (default: 60). Log filtered-out findings: "Auto-dismissed N low-confidence findings (below threshold [T]). Use discovery-confidence-threshold: 0 to see all."

4.5 Group by Category

Group remaining findings by category for Phase 5 presentation.

4.6 Embedded Mode Exit

If in embedded mode (called from session-end): STOP HERE. Return structured findings to the caller using this schema:

Embedded mode return schema:

{
  "findings": [
    {"probe": "string", "category": "string", "severity": "critical|high|medium|low", "confidence": 0-100, "file": "string", "line": number, "description": "string", "recommendation": "string"}
  ],
  "stats": {
    "probes_run": number,
    "findings_raw": number,
    "findings_verified": number,
    "false_positives": number,
    "user_dismissed": 0,
    "issues_created": 0,
    "by_category": {"<category>": {"findings": number, "actioned": 0}}
  }
}

Present both as structured data in your final output. Do not proceed to Phase 5.

Phase 5: Interactive Triage (Standalone Mode Only)

5.0 Load Triage State & Partition Findings

Before auto-defer and before presenting any findings for triage, load the persistent discovery triage state and filter findings through it:

  1. Call loadTriageState() from scripts/lib/discovery/triage-state.mjs (uses default path .orchestrator/metrics/discovery-triage.jsonl). Returns an empty Map if the file does not exist — no error.

  2. Call filterFindings({ findings: verifiedFindings, stateMap }) to partition findings into three buckets:

    • toShow — state is open, reopened, or no prior state entry (new findings — present for user triage)
    • suppressed — state is dismissed or accepted-as-known (skip silently)
    • tracked — state is promoted-to-#NNN (issue already filed; show as informational)
  3. Emit a one-line state banner before the summary table:

    Triage state: [N suppressed] suppressed (dismissed/accepted-as-known), [N tracked] tracked in existing issues. Presenting [N toShow] findings.
    

    Omit the banner entirely if all three counts are zero (first run).

  4. Render tracked findings as informational lines in the summary — NOT as interactive triage items:

    [INFO] Finding "<title>" (<file_path>) is tracked in #<issue_id> — not re-triaged.
    
  5. Continue Phase 5 triage using only toShow findings. The suppressed bucket requires no user interaction.

  6. After the user completes triage (Steps 1-4 below), append state changes to .orchestrator/metrics/discovery-triage.jsonl via appendTriageEntry() from triage-state.mjs:

    • User selects "Create issue" → append { fingerprint, state: 'promoted-to-#<issue_id>', issue_id: <N>, timestamp, session_id }
    • User selects "Dismiss -- intentional" or "Dismiss -- false positive" → append { fingerprint, state: 'dismissed', user_decision: '<reason>', timestamp, session_id }
    • User selects "Accept all" for batch → append one { fingerprint, state: 'open', ... } entry per finding (so they re-appear next run if not yet promoted)

5.1 Auto-Defer Low-Confidence Findings

Before presenting findings for triage, separate by confidence threshold:

  1. Findings with confidence >= threshold → present for interactive triage (below)
  2. Findings with confidence < threshold → auto-defer with summary: "Auto-deferred [N] low-confidence findings (score < [threshold]). Review with /discovery --include-deferred."
  3. List auto-deferred findings in a collapsed section (not interactive — informational only)

5.1 Present High-Confidence Findings

Present findings using AskUserQuestion -- NEVER plain text options. On Codex CLI where AskUserQuestion is unavailable, present as numbered Markdown lists.

Include confidence scores in the presentation:

[CRITICAL] (confidence: 85) hardcoded-values: API key found in src/config.ts:42
[HIGH] (confidence: 72) security-basics: eval() usage in src/utils/parser.ts:18
[MEDIUM] (confidence: 61) orphaned-annotations: TODO without issue in src/lib/auth.ts:55

Step 1: Summary

Present a findings overview table:

## Discovery Results

Probes run: [N] | Findings verified: [N] | False positives discarded: [N]

| Category | Critical | High | Medium | Low | Total |
|----------|----------|------|--------|-----|-------|
| Code     | ...      | ...  | ...    | ... | ...   |
| Infra    | ...      | ...  | ...    | ... | ...   |
| UI       | ...      | ...  | ...    | ... | ...   |
| Arch     | ...      | ...  | ...    | ... | ...   |
| Session  | ...      | ...  | ...    | ... | ...   |
| Audit    | ...      | ...  | ...    | ... | ...   |
| Vault    | ...      | ...  | ...    | ... | ...   |
| Feature  | ...      | ...  | ...    | ... | ...   |

Step 2: Critical + High Findings -- Review Individually

For each Critical or High finding, use AskUserQuestion (on Codex CLI where AskUserQuestion is unavailable, present as numbered Markdown lists):

AskUserQuestion({
  questions: [{
    question: "<severity> finding in <file_path> — what should happen with it?",
    header: "Finding",
    options: [
      { label: "Create issue (<severity>)", description: "Files it as priority::<severity>, so it is tracked outside this session. The code below is copied into the issue body.",
        preview: "<finding title>\n\n<file_path>:<line_number>\n```\n<matched_text with +/-3 lines context>\n```\n\n<description>\n\nRecommended fix: <recommended_fix>" },
      { label: "Adjust priority", description: "Same issue, a priority you pick — this question then comes back with the new label." },
      { label: "Dismiss -- intentional", description: "The code is deliberate. Nothing is filed, and the finding stays only in this run's report." },
      { label: "Dismiss -- false positive", description: "The probe misread the code. Nothing is filed; worth reporting if the same probe misfires again." }
    ],
    multiSelect: false
  }]
})

If user selects "Adjust priority", ask which priority with another AskUserQuestion. On Codex CLI where AskUserQuestion is unavailable, present as numbered Markdown lists.

Step 3: Medium + Low Findings -- Review Batched

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
50
Forks
7
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
discovery-kanevry
Source
github.com/kanevry/session-orchestrator