evolve
SkillMonitoring & opsUse this skill when extracting session patterns into reusable learnings. Three modes: analyze (extract from session history), review (edit/manage existing learnings), list (display active learnings). Manages .orchestrator/metrics/learnings.jsonl.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the evolve skill
What this skill tells your AI
The instructions your AI receives, as published by kanevry/session-orchestrator in skills/evolve/SKILL.md and read by ahel’s review.
Platform Note: State files use the platform's native directory:
.claude/(Claude Code),.codex/(Codex CLI), or.cursor/(Cursor IDE). Shared metrics live in.orchestrator/metrics/(v2) with fallback to<state-dir>/metrics/for pre-v2.0 legacy data. Seeskills/_shared/platform-tools.md.
Evolve Skill
Phase 0: Bootstrap Gate
Read skills/_shared/bootstrap-gate.md and execute the gate check. If the gate is CLOSED, invoke skills/bootstrap/SKILL.md and wait for completion before proceeding. If the gate is OPEN, continue to Phase 1.
Phase 1: Config & Data Loading
Telemetry start marker (#1200): note the current wall-clock time before Step 1.1 runs (e.g. date +%s%3N, or the coordinator's own turn-start instant). Every orchestrator.evolve.completed emit in Phase 1 / Phase 3 below reports duration_ms (placeholder DURATION_MS) as the elapsed milliseconds since this marker — same in-memory-value convention as CT/AC/ASK/DROP in skills/session-end/SKILL.md's orchestrator.handover.gated emits.
1.1 Read Session Config
Read and parse Session Config per skills/_shared/config-reading.md. Store result as $CONFIG.
1.2 Check Persistence
Extract persistence from $CONFIG. If persistence is false, abort with message:
"Learnings require persistence to be enabled in Session Config. Add
persistence: trueto your Session Config block (CLAUDE.md for Claude Code, AGENTS.md for Codex CLI)."
Telemetry on abort (#1200, #1206): before stopping, emit the abort form of the run-completion event. Kept as a minimal emit-event.mjs call, not routed through scripts/sweep-expired-learnings.mjs — no store-write CLI has run yet at this gate (it fires before Step 1.4 even reads learnings.jsonl), so there is no mechanical pipeline call site to fold this emit into, unlike the Step 3.5(5)/(6) success path below:
node scripts/emit-event.mjs --type orchestrator.evolve.completed --payload \
"$(node -e "process.stdout.write(JSON.stringify({aborted: 'persistence-disabled', reason: 'Learnings require persistence to be enabled in Session Config.'.slice(0,300), duration_ms: DURATION_MS}))")"
1.3 Determine Mode
Read mode from $ARGUMENTS:
- If empty or not provided, default to
analyze - Valid modes:
analyze,review,list,dialectic - If invalid mode provided, report error and list valid modes
1.4 Load Data
Lazy-create defensive (#185): If .orchestrator/metrics/learnings.jsonl does not exist (pre-#185 repo or bootstrap skipped), create an empty file and emit an info log — do NOT hard-fail:
LEARNINGS_FILE=".orchestrator/metrics/learnings.jsonl"
if [[ ! -f "$LEARNINGS_FILE" ]]; then
mkdir -p "$(dirname "$LEARNINGS_FILE")"
: > "$LEARNINGS_FILE"
echo "info(#185): auto-created $LEARNINGS_FILE (was missing)" >&2
fi
This defensive step is idempotent and cheap — it ensures /evolve analyze|review|list never fails because of a missing artifact file.
- Read
.orchestrator/metrics/sessions.jsonl(session history). If it does not exist, check<state-dir>/metrics/sessions.jsonlas a legacy fallback (where<state-dir>is.claude/,.codex/, or.cursor/per platform). If neither exists, warn: "No session history found. Run at least one session first." - Read
.orchestrator/metrics/learnings.jsonlif it exists. If not found, check<state-dir>/metrics/learnings.jsonlas a legacy fallback. - Count existing learnings, note any where
expires_at< current date (expired)
Phase 2: Mode Dispatch
Route based on mode:
analyze→ Phase 3review→ Phase 4list→ Phase 5dialectic→ Phase 6
Phase 3: Analyze Mode (default)
Extract learnings from session history.
Vault Integration: If
vault-integration.enabledistruein Session Config, confirmed learnings are mirrored to the configured Obsidian vault after the atomic write (Step 3.5, step 9). Seedocs/session-config-reference.mdfor thevault-integrationconfig block.
Step 3.1: Read Session Data
-
Read all entries from
.orchestrator/metrics/sessions.jsonl(or<state-dir>/metrics/sessions.jsonlif the v2 path does not exist — see Phase 1.4 fallback) -
Parse each JSONL line as JSON
-
Sort by
completed_atdescending (most recent first) -
If no sessions found, abort: "No session data available. Complete at least one session before running evolve." Telemetry on abort (#1200, #1206): before stopping, emit — same minimal
emit-event.mjscall as Phase 1.2's abort, and for the same reason: this gate fires before the Step 3.5(5)sweep-expired-learnings.mjs --prunecall exists to fold the emit into:node scripts/emit-event.mjs --type orchestrator.evolve.completed --payload \ "$(node -e "process.stdout.write(JSON.stringify({aborted: 'no-session-data', reason: 'No session data available. Complete at least one session before running evolve.'.slice(0,300), duration_ms: DURATION_MS}))")"
Step 3.1b: Read Extra Sources (#638)
When evolve.extra-sources is configured in Session Config (default [] ⇒ this step is a no-op), /evolve consumes OUT-OF-BAND domain measurement sidecars to surface domain-regression learnings.
READ-ONLY contract: /evolve NEVER runs the domain measurement. The measurement (e.g. an eval-learn regression harness) runs elsewhere and writes a sidecar JSON; this step only READS that sidecar's output. Never shell out to produce the sidecar from here.
For each configured extra-sources entry {path, kind, learning-type}:
- Read the sidecar at
path(parser-validated as repo-relative, with absolute paths and..escape segments dropped before this step, then resolved against the repo root). If the file is missing or unreadable, skip with a WARN (evolve: extra-source not found: <path>) — do not abort the whole run. - Schema-gate the sidecar against the
kind's expected shape. Forkind: regression-flagsthe schema is{ flags: [ { metric, baseline, recent, delta } ] }. If the parsed JSON does not match (missingflagsarray, or a flag missing a required field), skip with a WARN (evolve: extra-source <path> failed regression-flags schema gate) — never guess at a different shape. - Emit one
domain-regressionlearning candidate per flag that is PERSISTENT — i.e. the samemetricregressed across ≥2 consecutive sessions (cross-reference prior sessions' sidecar reads or the existing learnings store for the samesubject). A one-off flag is noise; only a persistent regression earns a candidate.type:learning-typefrom the entry (registered enum valuedomain-regression)subject: the flag'smetricinsight: a human-readable regression statement (e.g. "metric<metric>regressed: baseline → recent (delta ) persisting across ≥2 sessions")evidence:baseline → recent(the concrete data points from the sidecar)confidence/expires_at: derived via the existing confidence + decay infrastructure (Step 3.5), exactly as for the built-in learning types.domain-regressioncarries a 60-day TTL (LEARNING_TTL_DAYS).
- Candidates flow into the SAME Step 3.4 AskUserQuestion confirmation + Step 3.5 write path as the built-in learning types — there is no separate write path.
Step 3.2: Pattern Extraction
For each of the 9 built-in analyzer learning types, apply these heuristics:
1. fragile-file (type: fragile-file)
- Look at wave data: if the same file appears in 3+ waves'
files_changedwithin a session, it is fragile - Cross-session: if a file appears in 3+ different sessions'
files_changed, flag it - Subject = file path (relative to project root)
2. effective-sizing (type: effective-sizing)
- Compare
total_agentsandtotal_wavesacross session types - Calculate average agents per wave for each session type
- Subject = canonical identifier like
deep-session-sizingorfeature-session-sizing - Insight = "Deep sessions average X agents across Y waves" or "Feature sessions work well with X agents/wave"
- Over-delivery ratio aggregation (#730/H4, #794.7): compute the MEDIAN of
waves[].over_delivery_ratioacross the last ~5sessions.jsonlrecords of the samesession_type, filtered to waves whoseroleis notDiscovery/Finalizationand which carry the field (skip records lacking the field — pre-#730; also skip Discovery/Finalization waves, whose planned set is empty by design). This exclusion clause is intentionally identical toskills/session-plan/SKILL.mdStep 0.5 "Over-delivery sizing" — keep the two wordings in sync on edit. Fold the median into this candidate'sinsight/evidencefields — e.g.evidence:"median_over_delivery_ratio: 1.4 (n=12 waves, session_type=deep)"— sosession-planStep 0.5 can read the ratio from theeffective-sizinglearning first, falling back to its own directsessions.jsonlscan only when no such learning exists.
3. recurring-issue (type: recurring-issue)
- Look at
agent_summary— iffailedorpartial> 0 across multiple sessions, flag - Check wave
qualityfields — repeated failures indicate recurring issues - Subject = issue pattern identifier (e.g., "test-failures-in-wave-execution", "lint-regressions")
4. scope-guidance (type: scope-guidance)
- Cross-reference
effectiveness.planned_issuesvseffectiveness.completion_rate - Skip sessions that lack the
effectivenessfield (early sessions may not have it) - If completion_rate is consistently 1.0 with N issues, note "N issues per session works well"
- If completion_rate < 0.7, note "scope was too large"
- Subject =
optimal-scope-per-session-type
5. deviation-pattern (type: deviation-pattern)
Ownership Reference: See
skills/_shared/state-ownership.md. evolve has read-only access to STATE.md.
- Read
<state-dir>/STATE.mdif it exists and check## Deviationssection - Cross-reference with session duration vs planned waves
- Subject = pattern name (e.g., "scope-creep-in-feature-sessions", "underestimated-complexity")
6. stagnation-class-frequency (type: stagnation-class-frequency)
- Read
stagnation_eventsfrom the most recent 5 sessions insessions.jsonl(skip sessions lacking the field — they predate #84). - For each
(file, error_class)pair appearing in ≥2 sessions, extract a candidate:- Subject =
<file>:<error_class>(e.g.,skills/wave-executor/wave-loop.md:edit-format-friction) - Insight = "File has <error_class> stagnation in recent sessions — candidate for pre-edit grounding (#85)."
- Evidence = " sessions with stagnation_events for this file/class"
- Subject =
- These learnings feed #85 (pre-edit grounding injection) when it ships — high-frequency pairs trigger grounding.
7. hardware-pattern (type: hardware-pattern)
v3.1.0 / Sub-Epic #160 (C2, issue #171). Keyed on
host_classrather than project — surfaces hardware-bound problems that affect the user across every repo on the same machine. Complements the project-keyed types above.
- Read
.orchestrator/metrics/events.jsonl(session + wave events) and the registrysweep.logat~/.config/session-orchestrator/sessions/sweep.log. Both are optional — missing files produce no candidates. - Invoke
scripts/lib/hardware-pattern-detector.mjs→detectHardwarePatterns({events, sweepLogEntries, thresholds}). Thresholds come from Session Configresource-thresholdswhen present, falling back toDEFAULT_THRESHOLDS. - Five detection signals (aggregated per
(signal, host_class)pair, ≥2 occurrences required):- oom-kill —
orchestrator.turn.stopped(or its deprecated aliasorchestrator.session.stopped, whichhooks/on-stop.mjsstill emits withdeprecated: trueuntil 2027-03-06) withexit_code: 137or OOM-marker inerror. Both names are accepted for the deprecation window because every OOM record already on disk carries only the legacy name; the detector's set lives inOOM_TERMINAL_EVENTS(scripts/lib/hardware-pattern-detector.mjs) and drops the alias on that date. - heartbeat-gap — registry sweep-log entries with
gap_minutesaboveresource-thresholds.zombie-threshold-min - concurrent-session-pressure — session-start events with
peer_count ≥ concurrent-sessions-warn - disk-full — events whose
errormatchesENOSPC/ "no space left" - thermal-throttle — events whose
resource_snapshot.cpu_load_pctcrossescpu-load-max-pct
- oom-kill —
- Each candidate is piped through
candidateToLearning()→validateLearning(). Defaultscopeisprivate(in-repo only). To promote topublic, the user runsnpm run share:hw-learnings -- --promote(C3 export). This anonymizes eachprivatehardware-pattern entry, validates via the privacy contract, and appends apublictwin tolearnings.jsonl(original preserved). Use--dry-runto preview without writing. - Subject convention:
<signal>::<host_class>(e.g.,oom-kill::macos-arm64-m3pro). The::separator avoids colliding with project-keyed subjects. - Confidence starts at 0.5 like other learning types, but decay is slower in practice: hardware stays the same longer than code. This is an emergent property of the existing expire-after-N-days policy applied to a mostly-stable
host_class— no special-casing needed. - Presentation in step 3.5 (see below): render hardware-patterns in a dedicated section titled
## Hardware Patterns (keyed on host_class)after the project-keyed patterns. This makes the source of the learning obvious to the user at confirmation time.
8. autopilot-effectiveness (type: autopilot-effectiveness)
v3.2 Autopilot / Sub-Epic #271 (issue #298). Compares manual vs. autopilot session outcomes per mode (housekeeping, feature, deep) so the loop can learn whether walk-away runs preserve quality. Complements the project-keyed and hardware-keyed types above.
- Read
.orchestrator/metrics/autopilot.jsonl(one record per autopilot loop run) and.orchestrator/metrics/sessions.jsonl(manual + autopilot session outcomes). Both are optional — missing files produce no candidates. - Invoke
scripts/lib/evolve/autopilot-effectiveness.mjs→analyze(autopilotRuns, sessions). The module pairs records bymodeand compares completion-rate, carryover-rate, kill-switch frequency, and quality-gate pass-rate between the two populations. - Data-gating contract: the analyzer requires ≥20 paired manual+autopilot runs per mode before emitting any candidates. Below that threshold the function returns
[](empty input contract) — evolve simply skips this type for that mode and reports nothing. This prevents premature conclusions from small samples (#297 calibration depends on the same threshold). - Subject convention:
<mode>-manual-vs-autopilot(e.g.,housekeeping-manual-vs-autopilot,feature-manual-vs-autopilot,deep-manual-vs-autopilot). One subject per mode that crosses threshold. - Insight = "Autopilot sessions complete at % vs. manual % (Δ pp across N pairs)" or analogous carryover/kill-switch framing when those signals dominate.
- Confidence starts at 0.5 like other learning types; lifecycle ±0.15 / -0.20 via the existing dedupe-and-update infrastructure in Step 3.3 — no special-casing.
- Each candidate is piped through
candidateToLearning()→validateLearning()exactly like the other types. Defaultscopeisprivate(autopilot RUN data is per-host until the user opts in to share). (refs #298)
9. autonomy-verdict (type: autonomy-verdict)
Dispatcher Autonomy / P3.5 (issue #683). Synthesizes per-repo or per-scope autonomy readiness from autopilot run outcomes plus advisory skill-judge signals. Complements
autopilot-effectiveness: type 8 asks whether autopilot preserves quality by mode; this type asks whether a repo/scope is ready for more dispatcher autonomy.
- Read
.orchestrator/metrics/autopilot.jsonl,.orchestrator/metrics/sessions.jsonl, and.orchestrator/metrics/skill-judgments.jsonl. All are optional — missing files produce no candidates. - Invoke
scripts/lib/evolve/autonomy-verdict.mjs→analyze(autopilotRuns, sessions, skillJudgments, { repo | scope }). The analyzer reuses the type-8 mode rollups and combines them with counted skill-judgeapplied/completedsignals. - Data-gating contract: the analyzer requires ≥1 autopilot run and ≥1 canonical advisory skill-judge judgment (
schema_version: 1,event: "judged",advisory: true) before emitting a candidate. Below that threshold it returns[]so/evolve analyzestays quiet during cold-start. - Subject convention:
<repo-or-scope>-autonomy-readiness(e.g.,session-orchestrator-autonomy-readiness). - Insight frames the readiness verdict (
ready,watch, ornot-ready), the combined score, and the signal counts. Evidence includes the normalized scope, verdict, autopilot summary, and skill-judge summary. - Confidence is derived in the analyzer from signal volume, judge confidence, and score separation, then flows through the existing dedupe-and-update infrastructure in Step 3.3. Default
scopeisprivatebecause autopilot and skill-judge data are host/session-local. (refs #683)
Step 3.2b: Zero Patterns Check
If no patterns were extracted across all built-in analyzers and configured extra sources, report: "No patterns found in session history. This can happen with very few sessions or sessions that lack detailed wave/agent data." and skip to end (do not proceed to AskUserQuestion).
Step 3.3: Deduplicate Against Existing Learnings
For each extracted pattern, check if a learning with same type + subject already exists in learnings.jsonl:
- If exists: propose confidence update (+0.15 if confirmed by new evidence, -0.2 if contradicted)
- If new: propose as new learning with confidence 0.5
This match is exact string equality on type + subject — it is blind to two records that say the same thing in different words, and it cannot detect a contradiction at all. The -0.2 if contradicted branch above has therefore had no producer since it was written. Step 3.3b is that producer.
Step 3.3b: Relation Judgment (#1016)
Cadence: once per candidate. Step 3.2b's zero-patterns check and Step 3.4's single AUQ are once-per-run; Step 3.5's write is once-per-run. This step is the only per-candidate one in Phase 3 — the pool build happens once, the judgment runs for each pattern that seeds a pool.
Runs in
/evolve, never in a wave. The pool build is O(N²) over the candidate + corpus union (~13 ms at N=100 records; the viability boundary is ~N=2000)./evolveis operator-invoked and off the dispatch hot path — that is the whole reason this lives here and not inskills/wave-executor/. Do not invoke it from a wave prompt, an inter-wave checkpoint, or a hook.
Skip this step entirely when .orchestrator/metrics/learnings.jsonl is absent or holds fewer than 2 entries — with no corpus there is no relation to judge.
-
Pool. Call
buildCandidatePools(records, { now })fromscripts/lib/learnings/candidates.mjs, passing the union of this run's extracted candidates and the on-disk corpus. It returns{pools, duplicates, stats}:duplicatesare the exact-learning_keygroups (already certain — no judgment needed), and eachpools[]entry is{seed, candidates}wherecandidates[].recordis a bounded, per-seed, non-transitive neighbour set. No clustering, no transitive closure: a neighbour of a neighbour is not a neighbour. -
Judge, per candidate that seeds a pool.
buildJudgmentInput({candidate, neighbours})thenjudgeCandidate(input, { judge }), both fromscripts/lib/learnings/judgment.mjs.buildJudgmentInputreturnsnullfor a candidate with no usableid— skip that candidate, do not judge it.judgeis the injected verdict provider: on Claude Code the coordinator reads theinputenvelope and returns the JSON object itsoutput_contractfield describes. There is no subagent type for this — do not dispatch one (#614: a read-only agent that must write its own sidecar never fires). -
Apply, through the one choke point.
applyVerdict(verdict, effects)is the only place a judgment may become an effect. In/evolveevery effect handler is a proposal recorder, never a writer:refine/supersede/mergerecord a proposed change, andproposeContradictionrecords a contradiction pair.applyVerdictresolves all four handlers before invoking any of them, so an unwired handler refuses the whole batch rather than applying the decisions that happened to come first. -
Fail closed.
verdict.ok === false(any of the eight failure modes — unparseable, partial, phantom_id, self_reference, empty, timeout, enum_violation, duplicate_target) means no relation was read, not "no relation exists". The candidate keeps its Step 3.3 exact-match verdict and nothing about it is surfaced as a relation. Never fall back to a default decision, never repair-retry a malformed verdict, and never render an unreadable judgment to the operator — surfacing a relation IS the claim, so a voided judgment must not reach the AUQ at all. -
Route into the existing gate. Every surviving decision becomes an OPTION in Step 3.4's AskUserQuestion, never an action:
contradict→ a contradiction pair, presented as its own category beside "duplicate". If the operator selects it, it feeds the-0.2 if contradictedbranch in Step 3.3 above, applied by Step 3.5(3) — which deliberately does NOT resetexpires_at.supersede/merge→ an omit-the-loser (or replace-both-with-one) proposal. If selected, the operator's next generation simply omits those ids and Step 3.5(5) archives them — never a hand-delete. The merged record must carry both sources' provenance in its ownevidence.refine→ an edit proposal against the existing record'sinsight/evidence.skip/abstain→ nothing is surfaced.
The brandmauer holds here, unchanged (#693 FA2/FA3). The judgment computes; it never writes. Every .claude/rules/ write and every learnings.jsonl write stays behind the operator's Step 3.4 selection and Step 3.5's --prune invocation.
Named ceiling (revisit trigger). A supersede or merge executed through Step 3.5(5) is tagged _archive_reason: "superseded" with a _superseded_by tombstone only when the two records share type + non-empty subject — that is pruneLearnings()'s own consolidation pass. A cross-wording pair (the exact case this step exists to find) does not share a subject, so its loser is archived pruned instead: still in the corpus, still resolvable by id, but the archive record does not name its replacement. Revisit when the CLI grows per-record drop routing, or when an archive audit needs to answer "what replaced this?" for cross-wording merges.
Step 3.4: Present Findings via AskUserQuestion
Present extracted patterns to the user for confirmation. Use AskUserQuestion with multiSelect: true:
On Codex CLI where AskUserQuestion is unavailable, present as a numbered Markdown list.
AskUserQuestion({
questions: [{
question: "Which of the patterns extracted from this session's history should be saved?",
header: "Speichern?",
options: [
{
label: "[type] subject",
description: "insight | evidence: ... | confidence: 0.5 (new) or +0.15 (update)"
},
...
{
label: "Skip all",
description: "Do not save any learnings this time"
}
],
multiSelect: true
}]
})
If user selects "Skip all" or selects nothing, abort gracefully: "No learnings saved."
Step 3.5: Write Confirmed Learnings
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 50
- Forks
- 7
- Last commit
- Sep 2026
- Hacker News mentions
- 20
Advanced
- Catalog kind
- skill
- Gateway key
evolve- Source
- github.com/kanevry/session-orchestrator