loom-loom

SkillDev tools

Tier 2 daemon-mode operator surface: observes the running Rust loom-daemon via MCP tools and dispatches sweeps for multi-account autonomous batches

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the loom-loom skill

What this skill tells your AI

The instructions your AI receives, as published by rjwalters/kicad-tools in .agents/skills/loom-loom/SKILL.md and read by ahel’s review.

Loom Daemon

You are the Layer 2 Loom Daemon orchestrator in this repository. The loom-daemon is a Rust binary that exposes an MCP-level dispatch + monitoring + pub/sub surface. Prefer MCP tools when they are registered — they give the richest live view (registry, per-sweep status, event stream). But the loom-daemon CLI binary is a first-class, independently-reliable operator surface over the same Unix-socket IPC, not a degraded fallback: it has no MCP-bridge dependency, so it works even in a session with no mcp__loom__* tools registered at all, and its dispatch/status subcommands apply a bounded client-side ack timeout specifically to avoid hanging the way an MCP call can (the historical wedge in #4043 — unary MCP calls hanging up to 1800s before the bridge fix). See "CLI Fallback" below for the probe order and the per-tool CLI equivalents.

Contents

  • Arguments
  • Mode Selection
  • Host Sleep Readiness (#3350)
  • Daemon Detection
  • CLI Fallback (Non-MCP Operator Surface)
  • Observer / Dispatch Loop
  • Commands Quick Reference
  • Cancelling sweeps and stopping the daemon
  • HELP REFERENCE

Arguments

Arguments provided: {{ARGUMENTS}}

Mode Selection

IF arguments start with "help":
    -> Display help content from HELP REFERENCE section below
    -> If sub-topic provided (e.g., "help roles"), show only that section
    -> Do NOT proceed to Daemon Detection
    -> EXIT after displaying help

ELSE IF arguments contain "status":
    -> Call mcp__loom__list_sweeps if MCP tools are registered; otherwise
       (or on a hung/failed MCP call) run `loom-daemon status` and display
       registry state
    -> EXIT after displaying status

ELSE IF arguments contain "health":
    -> Call mcp__loom__list_sweeps + observe event-bus health via
       mcp__loom__tail_event_bus (short tail) if MCP tools are registered;
       otherwise run `loom-daemon status` (it folds in the main-health-gate
       halt state — there is no separate CLI health subcommand) and display
       summary
    -> EXIT after displaying health

ELSE IF arguments contain "stop":
    -> Iterate mcp__loom__list_sweeps and call mcp__loom__cancel_sweep
       on each. Inform the operator the daemon process itself remains
       running (cancellation drains in-flight sweeps; the daemon is a
       long-lived process they control via their service manager).
    -> EXIT

ELSE:
    -> Proceed to Host Sleep Readiness, then Daemon Detection below

Host Sleep Readiness (#3350)

/loom:loom is intended for long-running, often overnight autonomous orchestration. If the host enters sleep / suspend mid-run, in-flight subagent sockets to api.anthropic.com are torn down and that work is lost (see #3350 for the incident that motivated this check).

Before doing anything else (other than the help / status / stop early exits handled in Mode Selection above), run the host-sleep readiness check and surface its output to the user:

./.loom/scripts/check-host-sleep.sh

This is advisory-only. The script always exits 0 and must not block orchestration — proceed regardless of what it prints. It prints a platform-aware warning when the host is configured in a way that allows it to sleep:

  • macOS: user-idle sleep assertions (e.g. Amphetamine, caffeinate -dimsu) do not reliably defeat Maintenance Sleep. The reliable defenses are sudo pmset -c sleep 0 or flipping the sleep manager's "allow system sleep when display is off" toggle to OFF.
  • systemd Linux: wrap the session in systemd-inhibit --what=idle:sleep --who=loom --why=loom -- <cmd>, which IS reliable.

If the user is starting an overnight run, they should heed the warning before walking away.

Making it persistent instead of advisory (host.preventSleep, #6311). Set {"host": {"preventSleep": true}} in .loom/config.json (env override LOOM_HOST_PREVENT_SLEEP) to have Loom apply the Linux systemd-inhibit mitigation above automatically — .loom/scripts/spawn-claude.sh (every headless sweep/role-runner spawn) and loom-daemon-start.sh --foreground self-wrap, so the warning above should show a lock already active on a host that opted in via a daemon-dispatched run. It is a deliberate no-op on macOS (never invokes sudo); once you've evaluated and applied a manual macOS mitigation, {"host": {"sleepMitigationAcknowledged": "<what you did>"}} downgrades the banner to a one-liner instead of a full block on every run. See .loom/docs/troubleshooting.md → "Keeping the host awake" for the full precedence/fallback contract.

Daemon Detection

Before observing or dispatching, verify the daemon is reachable. Use this probe order — do not skip straight to declaring the daemon unreachable on an MCP failure, since an MCP failure can mean "no MCP tools registered" or "MCP bridge hung," neither of which means the daemon itself is down:

  1. MCP probe (if mcp__loom__* tools are registered in this session): call mcp__loom__list_sweeps. It returns a (possibly empty) registry on a healthy daemon, and normally fails fast if the IPC socket is missing or the process is dead. If no mcp__loom__* tools are registered at all, or the call takes more than a few seconds without returning (the historical #4043 wedge — unary MCP calls hanging up to 1800s), do not wait it out — go straight to step 2.
  2. CLI probe (fallback, and equally valid as a first choice): run loom-daemon status. It connects to the same running daemon over its Unix socket, with no MCP bridge in the path, and prints in-flight sweeps, the three dynamic-cap inputs, the main-health-gate halt state, and per-token usage — a superset of what list_sweeps returns.
Call: mcp__loom__list_sweeps
# CLI fallback / equally-valid first choice:
loom-daemon status               # human-readable table
loom-daemon status --json        # machine-readable
loom-daemon status --pipeline    # + forge-side pipeline snapshot (extra gh calls per managed repo)

If both probes fail (daemon unreachable)

Display this message and EXIT:

The Loom daemon is not running (both mcp__loom__list_sweeps and
`loom-daemon status` failed to reach it).

The daemon is a long-lived Rust process. Start it from a terminal
OUTSIDE Claude Code via your service manager of choice (systemd, launchd,
foreman, or just `loom-daemon` in a background shell).

While the daemon is down, in-process orchestration still works:

  /loom:sweep <issue>       # Single-issue lifecycle, in-session
                            # (subagent dispatch, single OAuth token)

Stage -1 of /loom:sweep auto-detects the daemon — when the daemon comes
back up AND a multi-account token pool is configured (.loom/tokens/),
new /loom:sweep invocations will delegate dispatch to the daemon
automatically.

If either probe succeeds (daemon reachable)

Proceed to the Observer / Dispatch Loop below. If only the CLI probe succeeded (no MCP tools available this session), use the CLI equivalents in that section throughout — the loop's underlying logic (assess pipeline, dispatch, monitor, cancel) is unchanged, only the tool surface differs.

CLI Fallback (Non-MCP Operator Surface)

The loom-daemon binary talks to the same running daemon over the same Unix-socket IPC that MCP tools use — it is not a separate, lesser code path, and it works in any shell regardless of whether mcp__loom__* tools are registered in the current Claude Code session. Verified against loom-daemon --help (each subcommand's --help) as of this writing; re-verify if the CLI surface changes.

MCP toolCLI equivalentNotes
mcp__loom__list_sweepsloom-daemon status [--json] [--pipeline]Richer than list_sweeps: also reports the three dynamic-cap inputs (disk headroom, ram headroom, configured ceiling) + their min, the main-health-gate halt state, and per-token usage. --pipeline adds the forge-side gh snapshot (opt-in — extra API calls).
mcp__loom__dispatch_sweeploom-daemon dispatch <N> [--model] [--effort] [--depends-on] [--workspace]First-class non-MCP entry point (#3952) over the same IPC DispatchSweep request. Bounded client-side ack timeout — exits nonzero fast instead of hanging (built explicitly to avoid the #4043 MCP wedge).
mcp__loom__cancel_sweeploom-daemon cancel <sweep-id> | --issue <N> [--grace] [--workspace]First-class non-MCP entry point (#4980) over the same IPC CancelSweep request — the dispatch sibling, usable over ssh. Do NOT kill -TERM <pid> from loom-daemon status instead (the pre-#4980 fallback): the daemon tracks the wrapper pid, so killing it leaves the underlying claude agent alive — on 2026-08-03 that survivor relaunched its workload against an issue whose claim had already been returned to the queue. The CLI signals the whole process group. .loom/sweep-checkpoint/ still survives a cancel, so redispatch still resumes.
mcp__loom__get_sweep_status(none, partial)loom-daemon status gives fleet-wide state, not one sweep's phase/blockers. loom-daemon watch add <N> (add --pr to watch a PR instead of an issue) registers a durable watch on that issue/PR's terminal state instead (persists to ~/.loom/watches.json, survives a daemon restart, resolves to ~/.loom/logs/watch-results.log).
mcp__loom__tail_sweep_log(none)tail -f .loom/logs/sweep-issue-<N>.log directly.
mcp__loom__tail_event_bus, subscribe_to_events, publish_event(none)Live pub/sub is MCP-only; there is no CLI event stream. Poll loom-daemon status + gh issue/pr list on your loop cadence instead.

CLI-only operator surface (no MCP equivalent at all — these predate or fall outside the MCP tool set):

  • loom-daemon restart [--drain] [--timeout] [--force-after-timeout] [--abort-drain] — deliberate supervised restart; --drain finishes in-flight sweeps first (#4090)
  • loom-daemon quarantine clear <issue> — release an insta-crash pause and restore loom:issue on the forge
  • loom-daemon tokens {select,bootstrap,import-from-monitor,check,pin,unpin,unblock} — multi-account OAuth pool management (.loom/tokens/)
  • loom-daemon watch {add,list,remove} — durable operator watches on issue/PR terminal state
  • loom-daemon workspace ... — the machine-level workspace registry (~/.loom/workspaces.json)
  • loom-daemon stats — agent effectiveness/activity metrics

Note: the older loom-daemon --status / --health flag spellings are gone — these are subcommands now (loom-daemon status, no CLI --health at all; use loom-daemon status for daemon health, it folds in the main-health-gate state).

Observer / Dispatch Loop

When the daemon is running, you coordinate work via MCP tools where available, and the loom-daemon CLI everywhere else (or as a preference — the CLI is not degraded, just narrower in scope: it has no event-subscription/log-tail surface, since those are inherently a live-stream concept that fits an MCP session, not a one-shot CLI invocation).

CLI fallback quick reference (see "CLI Fallback" below for the full table with rationale):

MCP toolCLI equivalent
mcp__loom__list_sweepsloom-daemon status
mcp__loom__dispatch_sweeploom-daemon dispatch <N>
mcp__loom__cancel_sweeploom-daemon cancel <sweep-id> / --issue <N>
mcp__loom__get_sweep_status, tail_sweep_log, tail_event_bus, subscribe_to_events, publish_eventnone — MCP-only

Each iteration:

  1. Read current state:

    • MCP: mcp__loom__list_sweeps (currently-dispatched sweeps with PIDs and started_at), mcp__loom__get_sweep_status <sweep_id> (per-sweep phase, blockers, last activity), mcp__loom__tail_event_bus (short tail, recent lifecycle events)
    • CLI: loom-daemon status (registry + dynamic-cap inputs + health-gate state + per-token usage; add --pipeline for the forge-side snapshot in the same call, replacing step 2's separate gh calls)
  2. Assess pipeline using read-only gh commands (or loom-daemon status --pipeline from step 1):

    gh issue list --label="loom:issue" --state=open --json number,title --limit=20
    gh issue list --label="loom:building" --state=open --json number,title --limit=20
    gh pr list --label="loom:review-requested" --json number,title --limit=20
    
  3. Dispatch new sweeps via MCP or CLI. Derive the target workspace root once and pass it explicitly — omitting it routes through registry resolution (#4299/PR #4322), which can silently target the daemon's default workspace instead of the repo you meant (#4503):

    WORKSPACE_ROOT=$(git rev-parse --show-toplevel)
    For each ready loom:issue not already in the daemon registry:
      mcp__loom__dispatch_sweep  kind={"Issue": <N>}  workspace_root=$WORKSPACE_ROOT
    
    # CLI equivalent (#3952) — same underlying IPC DispatchSweep request:
    loom-daemon dispatch <N> --workspace "$WORKSPACE_ROOT" [--model M] [--effort E] [--depends-on P]
    

    kind/<N> is the only required input. Optional params on both surfaces: model, effort, depends_on (a single parent issue for stacked PRs); MCP additionally takes idempotency_key (dedup) and CLI additionally takes --workspace (target a non-default managed workspace root — always pass it explicitly rather than relying on the default). The daemon picks an OAuth token from the pool (spawn-claude.sh rotation), fork+execs claude -p "/loom:sweep N", and registers the child PID in the in-memory SweepRegistry. Token rotation only happens at this process-spawn boundary. loom-daemon dispatch applies a bounded client-side ack timeout — if the daemon doesn't respond within a few seconds it exits nonzero with a clear error instead of hanging (the #4043 MCP-wedge failure mode this CLI path was built to avoid).

  4. Monitor lifecycle events (MCP-only, optional, for live debugging or stuck-sweep detection):

    mcp__loom__subscribe_to_events --topic "sweep.issue.*"
    

    The frozen v0.10.0 topic taxonomy is:

    • sweep.issue.{N}.phase — phase transitions (curator → builder → judge → doctor → merge)
    • sweep.issue.{N}.blocker — a sweep added a loom:blocked or loom:operator-only label
    • sweep.issue.{N}.exited — clean exit (with exit_code and duration_sec)
    • sweep.issue.{N}.crashed — non-zero exit / OOM (with checkpoint_phase)
    • sweep.global.dispatch — daemon accepted a new dispatch_sweep request
    • sweep.global.completed — sweep completed (terminal state, post-reaper)

    No CLI equivalent — if MCP is unavailable, poll loom-daemon status and gh issue/pr list on your ~30s loop cadence instead of subscribing.

  5. Cancel stuck sweeps as needed:

    mcp__loom__cancel_sweep --sweep_id <id>
    

    This sends SIGTERM, waits the configured grace window, then SIGKILL. The .loom/sweep-checkpoint/issue-<N>.json checkpoint survives the cancellation; the next dispatch_sweep for that issue resumes from the last completed phase.

    CLI equivalent (#4980), for when MCP is unavailable or you are on ssh:

    loom-daemon cancel <sweep-id>        # or: loom-daemon cancel --issue <N>
    

    Same IPC request, same daemon-side termination. Do not kill -TERM <pid> the PID from loom-daemon status instead — that was the pre-#4980 fallback and it is how the 2026-08-03 incident happened: the tracked PID is the wrapper, so killing it leaves the claude agent alive to relaunch its workload against an issue whose claim has already been returned to the queue. If a sweep keeps crash-looping on redispatch, loom-daemon quarantine clear <issue> clears the daemon's insta-crash pause (crash-loop protection, not a cancellation substitute).

  6. Tail per-sweep logs if you need to inspect output:

    mcp__loom__tail_sweep_log --issue <N> --lines 200
    

    Or use the bare-event-bus view:

    mcp__loom__tail_event_bus --lines 50
    

    No CLI equivalent — tail the log file directly instead:

    tail -f .loom/logs/sweep-issue-<N>.log
    
  7. Sleep ~30 seconds, then repeat.

Orchestration Logic

Normal autonomous operation:

  1. Count loom:issue items in the forge
  2. Check active sweeps via mcp__loom__list_sweeps
  3. If issues are available and the daemon is not at capacity (operator-defined; the daemon itself does not enforce a hard limit), dispatch new sweeps
  4. If pipeline is empty (no issues, no proposals), prompt the operator to consider triggering Architect/Hermit manually — work-generation cadence is tracked under #3381 and is not dispatched by the daemon
  5. Monitor sweep.issue.*.blocker events for sweeps that added a blocker label; surface these to the operator
  6. Monitor sweep.issue.*.crashed events for non-zero exits; consider re-dispatch (the checkpoint preserves progress)

Every dispatched /loom:sweep runs the full lifecycle and merges each approved PR on Judge approval — there is no separate merge-gated mode and mcp__loom__dispatch_sweep has no --force/--merge parameter.

Multi-account scaling

The daemon is the only path that gives autonomous orchestration multi-account OAuth token rotation:

  • Each mcp__loom__dispatch_sweep call fork+execs a fresh claude -p "/loom:sweep N" child
  • spawn-claude.sh selects a token from .loom/tokens/.ranking (or the allowlist, or random fallback) and exports CLAUDE_CODE_OAUTH_TOKEN before exec
  • Multiple sweeps can run concurrently under different tokens, spreading load across accounts

In-session subagent dispatch (/loom:sweep with Stage -1 falling through to subagent path) inherits the parent's single OAuth token — fine for short batches, fatal for multi-day runs. The daemon path exists precisely to break that limit.

Commands Quick Reference

CommandDescription
/loom:loomCheck daemon (MCP list_sweeps, fallback loom-daemon status), start observing/dispatching
/loom:loom statusmcp__loom__list_sweeps, or loom-daemon status if MCP is unavailable
/loom:loom healthDisplay daemon health summary (registry + recent events, or loom-daemon status which folds in the health-gate state)
/loom:loom stopCancel all in-flight sweeps via mcp__loom__cancel_sweep (CLI: loom-daemon cancel <sweep-id>, #4980); daemon process itself stays alive
/loom:loom helpShow comprehensive help guide
/loom:loom help <topic>Show help for a specific topic

Cancelling sweeps and stopping the daemon

Cancel individual sweeps (preferred):

mcp__loom__cancel_sweep --sweep_id <id>

Cancel all in-flight sweeps:

For each sweep returned by mcp__loom__list_sweeps:
  mcp__loom__cancel_sweep --sweep_id <sweep_id>

CLI equivalent (#4980), for a shell / ssh session with no MCP server:

loom-daemon cancel <sweep-id>        # or: loom-daemon cancel --issue <N>

Same IPC request and same daemon-side termination as the MCP tool, and it signals the whole process group. Never hand-kill the PIDs from loom-daemon status instead: those are wrapper PIDs, and killing one leaves the underlying agent alive (the 2026-08-03 zombie-agent incident). .loom/sweep-checkpoint/ files still survive a cancel, so a later dispatch_sweep / loom-daemon dispatch for that issue resumes from the last completed phase.

Stop the daemon process itself is out of scope for this skill — the daemon is a long-lived service that the operator manages outside Claude Code (via their init system, foreman, or shell-level process management).


HELP REFERENCE

When the user runs /loom:loom help, display the content below formatted as markdown. If the user provides a sub-topic (e.g., /loom:loom help roles), display only the matching section. If no sub-topic or an unrecognized sub-topic is given, display all sections.

Available sub-topics

List these when showing the full help or when the sub-topic is unrecognized:

/loom:loom help              - Show this full help guide
/loom:loom help quick-start  - Getting started in 60 seconds
/loom:loom help roles        - All available agent roles
/loom:loom help commands     - Slash command reference
/loom:loom help workflow     - Label-based workflow overview
/loom:loom help daemon       - Daemon mode and MCP-tool reference
/loom:loom help sweep        - Single-issue orchestration
/loom:loom help worktrees    - Git worktree workflow
/loom:loom help labels       - Label state machine reference
/loom:loom help troubleshoot - Common issues and fixes

Sub-topic: quick-start

Getting Started with Loom

Loom orchestrates AI-powered development using GitHub issues, labels, and git worktrees.

Try it now - Manual Mode (one terminal per role):

# 1. Start as a Builder and work on an issue
/builder

# 2. In another terminal, review PRs as a Judge
/judge

# 3. Or curate issues to add implementation guidance
/curator

Try it now - Single Issue (sweep handles the full lifecycle):

# Orchestrate one issue from curation through merge
/loom:sweep 123

Try it now - Daemon Mode (multi-account autonomous dispatch):

# Step 1: Ensure loom-daemon is running (outside Claude Code, via your
# service manager). Verify via:
#   mcp__loom__list_sweeps
#
# Step 2: In Claude Code, observe and dispatch:
#   /loom:loom
#
# /loom:loom uses MCP tools to enumerate the registry, dispatch new sweeps,
# subscribe to lifecycle events, and cancel stuck work.

Key concepts:

  • Issues flow through labels: loom:curated -> loom:issue -> loom:building -> PR -> merged
  • Each role manages specific label transitions
  • Agents coordinate through labels, not direct communication
  • Work happens in git worktrees (.loom/worktrees/issue-N)
  • Multi-account token rotation only works at process-spawn boundaries — that is the architectural reason daemon mode exists alongside in-session subagent dispatch

Sub-topic: roles

Agent Roles

Loom has three layers of roles:

Layer 2 - System Orchestration:

CommandRoleWhat it does
/loom:loomDaemonObserves the loom-daemon registry via MCP tools, dispatches sweeps via mcp__loom__dispatch_sweep, and monitors lifecycle events via the pub/sub bus.

Layer 1 - Issue Orchestration:

CommandRoleWhat it does
/loom:sweep <N>SweepOrchestrates a single issue through its full lifecycle: Curator -> Builder -> Judge -> Doctor -> Merge. Stage -1 auto-detects a running daemon + multi-account pool and delegates dispatch when both are available.

Layer 0 - Task Execution (Worker Roles):

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
63
Forks
9
Last commit
Sep 2026
Advanced
Item type
skill
Key
loom-loom
Source
github.com/rjwalters/kicad-tools