evolve

SkillMonitoring & ops

Use this skill when extracting session patterns into reusable learnings. Three modes: analyze (extract from session history), review (edit/manage existing learnings), list (display active learnings). Manages .orchestrator/metrics/learnings.jsonl.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the evolve skill

What this skill tells your AI

The instructions your AI receives, as published by kanevry/session-orchestrator in skills/evolve/SKILL.md and read by ahel’s review.

Platform Note: State files use the platform's native directory: .claude/ (Claude Code), .codex/ (Codex CLI), or .cursor/ (Cursor IDE). Shared metrics live in .orchestrator/metrics/ (v2) with fallback to <state-dir>/metrics/ for pre-v2.0 legacy data. See skills/_shared/platform-tools.md.

Evolve Skill

Phase 0: Bootstrap Gate

Read skills/_shared/bootstrap-gate.md and execute the gate check. If the gate is CLOSED, invoke skills/bootstrap/SKILL.md and wait for completion before proceeding. If the gate is OPEN, continue to Phase 1.

Phase 1: Config & Data Loading

Telemetry start marker (#1200): note the current wall-clock time before Step 1.1 runs (e.g. date +%s%3N, or the coordinator's own turn-start instant). Every orchestrator.evolve.completed emit in Phase 1 / Phase 3 below reports duration_ms (placeholder DURATION_MS) as the elapsed milliseconds since this marker — same in-memory-value convention as CT/AC/ASK/DROP in skills/session-end/SKILL.md's orchestrator.handover.gated emits.

1.1 Read Session Config

Read and parse Session Config per skills/_shared/config-reading.md. Store result as $CONFIG.

1.2 Check Persistence

Extract persistence from $CONFIG. If persistence is false, abort with message:

"Learnings require persistence to be enabled in Session Config. Add persistence: true to your Session Config block (CLAUDE.md for Claude Code, AGENTS.md for Codex CLI)."

Telemetry on abort (#1200, #1206): before stopping, emit the abort form of the run-completion event. Kept as a minimal emit-event.mjs call, not routed through scripts/sweep-expired-learnings.mjs — no store-write CLI has run yet at this gate (it fires before Step 1.4 even reads learnings.jsonl), so there is no mechanical pipeline call site to fold this emit into, unlike the Step 3.5(5)/(6) success path below:

node scripts/emit-event.mjs --type orchestrator.evolve.completed --payload \
  "$(node -e "process.stdout.write(JSON.stringify({aborted: 'persistence-disabled', reason: 'Learnings require persistence to be enabled in Session Config.'.slice(0,300), duration_ms: DURATION_MS}))")"

1.3 Determine Mode

Read mode from $ARGUMENTS:

  • If empty or not provided, default to analyze
  • Valid modes: analyze, review, list, dialectic
  • If invalid mode provided, report error and list valid modes

1.4 Load Data

Lazy-create defensive (#185): If .orchestrator/metrics/learnings.jsonl does not exist (pre-#185 repo or bootstrap skipped), create an empty file and emit an info log — do NOT hard-fail:

LEARNINGS_FILE=".orchestrator/metrics/learnings.jsonl"
if [[ ! -f "$LEARNINGS_FILE" ]]; then
  mkdir -p "$(dirname "$LEARNINGS_FILE")"
  : > "$LEARNINGS_FILE"
  echo "info(#185): auto-created $LEARNINGS_FILE (was missing)" >&2
fi

This defensive step is idempotent and cheap — it ensures /evolve analyze|review|list never fails because of a missing artifact file.

  1. Read .orchestrator/metrics/sessions.jsonl (session history). If it does not exist, check <state-dir>/metrics/sessions.jsonl as a legacy fallback (where <state-dir> is .claude/, .codex/, or .cursor/ per platform). If neither exists, warn: "No session history found. Run at least one session first."
  2. Read .orchestrator/metrics/learnings.jsonl if it exists. If not found, check <state-dir>/metrics/learnings.jsonl as a legacy fallback.
  3. Count existing learnings, note any where expires_at < current date (expired)

Phase 2: Mode Dispatch

Route based on mode:

  • analyze → Phase 3
  • review → Phase 4
  • list → Phase 5
  • dialectic → Phase 6

Phase 3: Analyze Mode (default)

Extract learnings from session history.

Vault Integration: If vault-integration.enabled is true in Session Config, confirmed learnings are mirrored to the configured Obsidian vault after the atomic write (Step 3.5, step 9). See docs/session-config-reference.md for the vault-integration config block.

Step 3.1: Read Session Data

  • Read all entries from .orchestrator/metrics/sessions.jsonl (or <state-dir>/metrics/sessions.jsonl if the v2 path does not exist — see Phase 1.4 fallback)

  • Parse each JSONL line as JSON

  • Sort by completed_at descending (most recent first)

  • If no sessions found, abort: "No session data available. Complete at least one session before running evolve." Telemetry on abort (#1200, #1206): before stopping, emit — same minimal emit-event.mjs call as Phase 1.2's abort, and for the same reason: this gate fires before the Step 3.5(5) sweep-expired-learnings.mjs --prune call exists to fold the emit into:

    node scripts/emit-event.mjs --type orchestrator.evolve.completed --payload \
      "$(node -e "process.stdout.write(JSON.stringify({aborted: 'no-session-data', reason: 'No session data available. Complete at least one session before running evolve.'.slice(0,300), duration_ms: DURATION_MS}))")"
    

Step 3.1b: Read Extra Sources (#638)

When evolve.extra-sources is configured in Session Config (default [] ⇒ this step is a no-op), /evolve consumes OUT-OF-BAND domain measurement sidecars to surface domain-regression learnings.

READ-ONLY contract: /evolve NEVER runs the domain measurement. The measurement (e.g. an eval-learn regression harness) runs elsewhere and writes a sidecar JSON; this step only READS that sidecar's output. Never shell out to produce the sidecar from here.

For each configured extra-sources entry {path, kind, learning-type}:

  1. Read the sidecar at path (parser-validated as repo-relative, with absolute paths and .. escape segments dropped before this step, then resolved against the repo root). If the file is missing or unreadable, skip with a WARN (evolve: extra-source not found: <path>) — do not abort the whole run.
  2. Schema-gate the sidecar against the kind's expected shape. For kind: regression-flags the schema is { flags: [ { metric, baseline, recent, delta } ] }. If the parsed JSON does not match (missing flags array, or a flag missing a required field), skip with a WARN (evolve: extra-source <path> failed regression-flags schema gate) — never guess at a different shape.
  3. Emit one domain-regression learning candidate per flag that is PERSISTENT — i.e. the same metric regressed across ≥2 consecutive sessions (cross-reference prior sessions' sidecar reads or the existing learnings store for the same subject). A one-off flag is noise; only a persistent regression earns a candidate.
    • type: learning-type from the entry (registered enum value domain-regression)
    • subject: the flag's metric
    • insight: a human-readable regression statement (e.g. "metric <metric> regressed: baseline → recent (delta ) persisting across ≥2 sessions")
    • evidence: baseline → recent (the concrete data points from the sidecar)
    • confidence / expires_at: derived via the existing confidence + decay infrastructure (Step 3.5), exactly as for the built-in learning types. domain-regression carries a 60-day TTL (LEARNING_TTL_DAYS).
  4. Candidates flow into the SAME Step 3.4 AskUserQuestion confirmation + Step 3.5 write path as the built-in learning types — there is no separate write path.

Step 3.2: Pattern Extraction

For each of the 9 built-in analyzer learning types, apply these heuristics:

1. fragile-file (type: fragile-file)
  • Look at wave data: if the same file appears in 3+ waves' files_changed within a session, it is fragile
  • Cross-session: if a file appears in 3+ different sessions' files_changed, flag it
  • Subject = file path (relative to project root)
2. effective-sizing (type: effective-sizing)
  • Compare total_agents and total_waves across session types
  • Calculate average agents per wave for each session type
  • Subject = canonical identifier like deep-session-sizing or feature-session-sizing
  • Insight = "Deep sessions average X agents across Y waves" or "Feature sessions work well with X agents/wave"
  • Over-delivery ratio aggregation (#730/H4, #794.7): compute the MEDIAN of waves[].over_delivery_ratio across the last ~5 sessions.jsonl records of the same session_type, filtered to waves whose role is not Discovery/Finalization and which carry the field (skip records lacking the field — pre-#730; also skip Discovery/Finalization waves, whose planned set is empty by design). This exclusion clause is intentionally identical to skills/session-plan/SKILL.md Step 0.5 "Over-delivery sizing" — keep the two wordings in sync on edit. Fold the median into this candidate's insight/evidence fields — e.g. evidence: "median_over_delivery_ratio: 1.4 (n=12 waves, session_type=deep)" — so session-plan Step 0.5 can read the ratio from the effective-sizing learning first, falling back to its own direct sessions.jsonl scan only when no such learning exists.
3. recurring-issue (type: recurring-issue)
  • Look at agent_summary — if failed or partial > 0 across multiple sessions, flag
  • Check wave quality fields — repeated failures indicate recurring issues
  • Subject = issue pattern identifier (e.g., "test-failures-in-wave-execution", "lint-regressions")
4. scope-guidance (type: scope-guidance)
  • Cross-reference effectiveness.planned_issues vs effectiveness.completion_rate
  • Skip sessions that lack the effectiveness field (early sessions may not have it)
  • If completion_rate is consistently 1.0 with N issues, note "N issues per session works well"
  • If completion_rate < 0.7, note "scope was too large"
  • Subject = optimal-scope-per-session-type
5. deviation-pattern (type: deviation-pattern)

Ownership Reference: See skills/_shared/state-ownership.md. evolve has read-only access to STATE.md.

  • Read <state-dir>/STATE.md if it exists and check ## Deviations section
  • Cross-reference with session duration vs planned waves
  • Subject = pattern name (e.g., "scope-creep-in-feature-sessions", "underestimated-complexity")
6. stagnation-class-frequency (type: stagnation-class-frequency)
  • Read stagnation_events from the most recent 5 sessions in sessions.jsonl (skip sessions lacking the field — they predate #84).
  • For each (file, error_class) pair appearing in ≥2 sessions, extract a candidate:
    • Subject = <file>:<error_class> (e.g., skills/wave-executor/wave-loop.md:edit-format-friction)
    • Insight = "File has <error_class> stagnation in recent sessions — candidate for pre-edit grounding (#85)."
    • Evidence = " sessions with stagnation_events for this file/class"
  • These learnings feed #85 (pre-edit grounding injection) when it ships — high-frequency pairs trigger grounding.
7. hardware-pattern (type: hardware-pattern)

v3.1.0 / Sub-Epic #160 (C2, issue #171). Keyed on host_class rather than project — surfaces hardware-bound problems that affect the user across every repo on the same machine. Complements the project-keyed types above.

  • Read .orchestrator/metrics/events.jsonl (session + wave events) and the registry sweep.log at ~/.config/session-orchestrator/sessions/sweep.log. Both are optional — missing files produce no candidates.
  • Invoke scripts/lib/hardware-pattern-detector.mjsdetectHardwarePatterns({events, sweepLogEntries, thresholds}). Thresholds come from Session Config resource-thresholds when present, falling back to DEFAULT_THRESHOLDS.
  • Five detection signals (aggregated per (signal, host_class) pair, ≥2 occurrences required):
    • oom-killorchestrator.turn.stopped (or its deprecated alias orchestrator.session.stopped, which hooks/on-stop.mjs still emits with deprecated: true until 2027-03-06) with exit_code: 137 or OOM-marker in error. Both names are accepted for the deprecation window because every OOM record already on disk carries only the legacy name; the detector's set lives in OOM_TERMINAL_EVENTS (scripts/lib/hardware-pattern-detector.mjs) and drops the alias on that date.
    • heartbeat-gap — registry sweep-log entries with gap_minutes above resource-thresholds.zombie-threshold-min
    • concurrent-session-pressure — session-start events with peer_count ≥ concurrent-sessions-warn
    • disk-full — events whose error matches ENOSPC / "no space left"
    • thermal-throttle — events whose resource_snapshot.cpu_load_pct crosses cpu-load-max-pct
  • Each candidate is piped through candidateToLearning()validateLearning(). Default scope is private (in-repo only). To promote to public, the user runs npm run share:hw-learnings -- --promote (C3 export). This anonymizes each private hardware-pattern entry, validates via the privacy contract, and appends a public twin to learnings.jsonl (original preserved). Use --dry-run to preview without writing.
  • Subject convention: <signal>::<host_class> (e.g., oom-kill::macos-arm64-m3pro). The :: separator avoids colliding with project-keyed subjects.
  • Confidence starts at 0.5 like other learning types, but decay is slower in practice: hardware stays the same longer than code. This is an emergent property of the existing expire-after-N-days policy applied to a mostly-stable host_class — no special-casing needed.
  • Presentation in step 3.5 (see below): render hardware-patterns in a dedicated section titled ## Hardware Patterns (keyed on host_class) after the project-keyed patterns. This makes the source of the learning obvious to the user at confirmation time.
8. autopilot-effectiveness (type: autopilot-effectiveness)

v3.2 Autopilot / Sub-Epic #271 (issue #298). Compares manual vs. autopilot session outcomes per mode (housekeeping, feature, deep) so the loop can learn whether walk-away runs preserve quality. Complements the project-keyed and hardware-keyed types above.

  • Read .orchestrator/metrics/autopilot.jsonl (one record per autopilot loop run) and .orchestrator/metrics/sessions.jsonl (manual + autopilot session outcomes). Both are optional — missing files produce no candidates.
  • Invoke scripts/lib/evolve/autopilot-effectiveness.mjsanalyze(autopilotRuns, sessions). The module pairs records by mode and compares completion-rate, carryover-rate, kill-switch frequency, and quality-gate pass-rate between the two populations.
  • Data-gating contract: the analyzer requires ≥20 paired manual+autopilot runs per mode before emitting any candidates. Below that threshold the function returns [] (empty input contract) — evolve simply skips this type for that mode and reports nothing. This prevents premature conclusions from small samples (#297 calibration depends on the same threshold).
  • Subject convention: <mode>-manual-vs-autopilot (e.g., housekeeping-manual-vs-autopilot, feature-manual-vs-autopilot, deep-manual-vs-autopilot). One subject per mode that crosses threshold.
  • Insight = "Autopilot sessions complete at % vs. manual % (Δ pp across N pairs)" or analogous carryover/kill-switch framing when those signals dominate.
  • Confidence starts at 0.5 like other learning types; lifecycle ±0.15 / -0.20 via the existing dedupe-and-update infrastructure in Step 3.3 — no special-casing.
  • Each candidate is piped through candidateToLearning()validateLearning() exactly like the other types. Default scope is private (autopilot RUN data is per-host until the user opts in to share). (refs #298)
9. autonomy-verdict (type: autonomy-verdict)

Dispatcher Autonomy / P3.5 (issue #683). Synthesizes per-repo or per-scope autonomy readiness from autopilot run outcomes plus advisory skill-judge signals. Complements autopilot-effectiveness: type 8 asks whether autopilot preserves quality by mode; this type asks whether a repo/scope is ready for more dispatcher autonomy.

  • Read .orchestrator/metrics/autopilot.jsonl, .orchestrator/metrics/sessions.jsonl, and .orchestrator/metrics/skill-judgments.jsonl. All are optional — missing files produce no candidates.
  • Invoke scripts/lib/evolve/autonomy-verdict.mjsanalyze(autopilotRuns, sessions, skillJudgments, { repo | scope }). The analyzer reuses the type-8 mode rollups and combines them with counted skill-judge applied/completed signals.
  • Data-gating contract: the analyzer requires ≥1 autopilot run and ≥1 canonical advisory skill-judge judgment (schema_version: 1, event: "judged", advisory: true) before emitting a candidate. Below that threshold it returns [] so /evolve analyze stays quiet during cold-start.
  • Subject convention: <repo-or-scope>-autonomy-readiness (e.g., session-orchestrator-autonomy-readiness).
  • Insight frames the readiness verdict (ready, watch, or not-ready), the combined score, and the signal counts. Evidence includes the normalized scope, verdict, autopilot summary, and skill-judge summary.
  • Confidence is derived in the analyzer from signal volume, judge confidence, and score separation, then flows through the existing dedupe-and-update infrastructure in Step 3.3. Default scope is private because autopilot and skill-judge data are host/session-local. (refs #683)

Step 3.2b: Zero Patterns Check

If no patterns were extracted across all built-in analyzers and configured extra sources, report: "No patterns found in session history. This can happen with very few sessions or sessions that lack detailed wave/agent data." and skip to end (do not proceed to AskUserQuestion).

Step 3.3: Deduplicate Against Existing Learnings

For each extracted pattern, check if a learning with same type + subject already exists in learnings.jsonl:

  • If exists: propose confidence update (+0.15 if confirmed by new evidence, -0.2 if contradicted)
  • If new: propose as new learning with confidence 0.5

This match is exact string equality on type + subject — it is blind to two records that say the same thing in different words, and it cannot detect a contradiction at all. The -0.2 if contradicted branch above has therefore had no producer since it was written. Step 3.3b is that producer.

Step 3.3b: Relation Judgment (#1016)

Cadence: once per candidate. Step 3.2b's zero-patterns check and Step 3.4's single AUQ are once-per-run; Step 3.5's write is once-per-run. This step is the only per-candidate one in Phase 3 — the pool build happens once, the judgment runs for each pattern that seeds a pool.

Runs in /evolve, never in a wave. The pool build is O(N²) over the candidate + corpus union (~13 ms at N=100 records; the viability boundary is ~N=2000). /evolve is operator-invoked and off the dispatch hot path — that is the whole reason this lives here and not in skills/wave-executor/. Do not invoke it from a wave prompt, an inter-wave checkpoint, or a hook.

Skip this step entirely when .orchestrator/metrics/learnings.jsonl is absent or holds fewer than 2 entries — with no corpus there is no relation to judge.

  1. Pool. Call buildCandidatePools(records, { now }) from scripts/lib/learnings/candidates.mjs, passing the union of this run's extracted candidates and the on-disk corpus. It returns {pools, duplicates, stats}: duplicates are the exact-learning_key groups (already certain — no judgment needed), and each pools[] entry is {seed, candidates} where candidates[].record is a bounded, per-seed, non-transitive neighbour set. No clustering, no transitive closure: a neighbour of a neighbour is not a neighbour.

  2. Judge, per candidate that seeds a pool. buildJudgmentInput({candidate, neighbours}) then judgeCandidate(input, { judge }), both from scripts/lib/learnings/judgment.mjs. buildJudgmentInput returns null for a candidate with no usable id — skip that candidate, do not judge it. judge is the injected verdict provider: on Claude Code the coordinator reads the input envelope and returns the JSON object its output_contract field describes. There is no subagent type for this — do not dispatch one (#614: a read-only agent that must write its own sidecar never fires).

  3. Apply, through the one choke point. applyVerdict(verdict, effects) is the only place a judgment may become an effect. In /evolve every effect handler is a proposal recorder, never a writer: refine / supersede / merge record a proposed change, and proposeContradiction records a contradiction pair. applyVerdict resolves all four handlers before invoking any of them, so an unwired handler refuses the whole batch rather than applying the decisions that happened to come first.

  4. Fail closed. verdict.ok === false (any of the eight failure modes — unparseable, partial, phantom_id, self_reference, empty, timeout, enum_violation, duplicate_target) means no relation was read, not "no relation exists". The candidate keeps its Step 3.3 exact-match verdict and nothing about it is surfaced as a relation. Never fall back to a default decision, never repair-retry a malformed verdict, and never render an unreadable judgment to the operator — surfacing a relation IS the claim, so a voided judgment must not reach the AUQ at all.

  5. Route into the existing gate. Every surviving decision becomes an OPTION in Step 3.4's AskUserQuestion, never an action:

    • contradict → a contradiction pair, presented as its own category beside "duplicate". If the operator selects it, it feeds the -0.2 if contradicted branch in Step 3.3 above, applied by Step 3.5(3) — which deliberately does NOT reset expires_at.
    • supersede / merge → an omit-the-loser (or replace-both-with-one) proposal. If selected, the operator's next generation simply omits those ids and Step 3.5(5) archives them — never a hand-delete. The merged record must carry both sources' provenance in its own evidence.
    • refine → an edit proposal against the existing record's insight / evidence.
    • skip / abstain → nothing is surfaced.

The brandmauer holds here, unchanged (#693 FA2/FA3). The judgment computes; it never writes. Every .claude/rules/ write and every learnings.jsonl write stays behind the operator's Step 3.4 selection and Step 3.5's --prune invocation.

Named ceiling (revisit trigger). A supersede or merge executed through Step 3.5(5) is tagged _archive_reason: "superseded" with a _superseded_by tombstone only when the two records share type + non-empty subject — that is pruneLearnings()'s own consolidation pass. A cross-wording pair (the exact case this step exists to find) does not share a subject, so its loser is archived pruned instead: still in the corpus, still resolvable by id, but the archive record does not name its replacement. Revisit when the CLI grows per-record drop routing, or when an archive audit needs to answer "what replaced this?" for cross-wording merges.

Step 3.4: Present Findings via AskUserQuestion

Present extracted patterns to the user for confirmation. Use AskUserQuestion with multiSelect: true:

On Codex CLI where AskUserQuestion is unavailable, present as a numbered Markdown list.

AskUserQuestion({
  questions: [{
    question: "Which of the patterns extracted from this session's history should be saved?",
    header: "Speichern?",
    options: [
      {
        label: "[type] subject",
        description: "insight | evidence: ... | confidence: 0.5 (new) or +0.15 (update)"
      },
      ...
      {
        label: "Skip all",
        description: "Do not save any learnings this time"
      }
    ],
    multiSelect: true
  }]
})

If user selects "Skip all" or selects nothing, abort gracefully: "No learnings saved."

Step 3.5: Write Confirmed Learnings

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
50
Forks
7
Last commit
Sep 2026
Hacker News mentions
20
Advanced
Catalog kind
skill
Gateway key
evolve
Source
github.com/kanevry/session-orchestrator