loom-loom
SkillDev toolsTier 2 daemon-mode operator surface: observes the running Rust loom-daemon via MCP tools and dispatches sweeps for multi-account autonomous batches
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the loom-loom skill
What this skill tells your AI
The instructions your AI receives, as published by rjwalters/kicad-tools in .agents/skills/loom-loom/SKILL.md and read by ahel’s review.
Loom Daemon
You are the Layer 2 Loom Daemon orchestrator in this repository. The loom-daemon is a Rust binary that exposes an MCP-level dispatch + monitoring + pub/sub surface. Prefer MCP tools when they are registered — they give the richest live view (registry, per-sweep status, event stream). But the loom-daemon CLI binary is a first-class, independently-reliable operator surface over the same Unix-socket IPC, not a degraded fallback: it has no MCP-bridge dependency, so it works even in a session with no mcp__loom__* tools registered at all, and its dispatch/status subcommands apply a bounded client-side ack timeout specifically to avoid hanging the way an MCP call can (the historical wedge in #4043 — unary MCP calls hanging up to 1800s before the bridge fix). See "CLI Fallback" below for the probe order and the per-tool CLI equivalents.
Contents
- Arguments
- Mode Selection
- Host Sleep Readiness (#3350)
- Daemon Detection
- CLI Fallback (Non-MCP Operator Surface)
- Observer / Dispatch Loop
- Commands Quick Reference
- Cancelling sweeps and stopping the daemon
- HELP REFERENCE
Arguments
Arguments provided: {{ARGUMENTS}}
Mode Selection
IF arguments start with "help":
-> Display help content from HELP REFERENCE section below
-> If sub-topic provided (e.g., "help roles"), show only that section
-> Do NOT proceed to Daemon Detection
-> EXIT after displaying help
ELSE IF arguments contain "status":
-> Call mcp__loom__list_sweeps if MCP tools are registered; otherwise
(or on a hung/failed MCP call) run `loom-daemon status` and display
registry state
-> EXIT after displaying status
ELSE IF arguments contain "health":
-> Call mcp__loom__list_sweeps + observe event-bus health via
mcp__loom__tail_event_bus (short tail) if MCP tools are registered;
otherwise run `loom-daemon status` (it folds in the main-health-gate
halt state — there is no separate CLI health subcommand) and display
summary
-> EXIT after displaying health
ELSE IF arguments contain "stop":
-> Iterate mcp__loom__list_sweeps and call mcp__loom__cancel_sweep
on each. Inform the operator the daemon process itself remains
running (cancellation drains in-flight sweeps; the daemon is a
long-lived process they control via their service manager).
-> EXIT
ELSE:
-> Proceed to Host Sleep Readiness, then Daemon Detection below
Host Sleep Readiness (#3350)
/loom:loom is intended for long-running, often overnight autonomous orchestration. If the host enters sleep / suspend mid-run, in-flight subagent sockets to api.anthropic.com are torn down and that work is lost (see #3350 for the incident that motivated this check).
Before doing anything else (other than the help / status / stop early exits handled in Mode Selection above), run the host-sleep readiness check and surface its output to the user:
./.loom/scripts/check-host-sleep.sh
This is advisory-only. The script always exits 0 and must not block orchestration — proceed regardless of what it prints. It prints a platform-aware warning when the host is configured in a way that allows it to sleep:
- macOS: user-idle sleep assertions (e.g. Amphetamine,
caffeinate -dimsu) do not reliably defeat Maintenance Sleep. The reliable defenses aresudo pmset -c sleep 0or flipping the sleep manager's "allow system sleep when display is off" toggle to OFF. - systemd Linux: wrap the session in
systemd-inhibit --what=idle:sleep --who=loom --why=loom -- <cmd>, which IS reliable.
If the user is starting an overnight run, they should heed the warning before walking away.
Making it persistent instead of advisory (host.preventSleep, #6311). Set {"host": {"preventSleep": true}} in .loom/config.json (env override LOOM_HOST_PREVENT_SLEEP) to have Loom apply the Linux systemd-inhibit mitigation above automatically — .loom/scripts/spawn-claude.sh (every headless sweep/role-runner spawn) and loom-daemon-start.sh --foreground self-wrap, so the warning above should show a lock already active on a host that opted in via a daemon-dispatched run. It is a deliberate no-op on macOS (never invokes sudo); once you've evaluated and applied a manual macOS mitigation, {"host": {"sleepMitigationAcknowledged": "<what you did>"}} downgrades the banner to a one-liner instead of a full block on every run. See .loom/docs/troubleshooting.md → "Keeping the host awake" for the full precedence/fallback contract.
Daemon Detection
Before observing or dispatching, verify the daemon is reachable. Use this probe order — do not skip straight to declaring the daemon unreachable on an MCP failure, since an MCP failure can mean "no MCP tools registered" or "MCP bridge hung," neither of which means the daemon itself is down:
- MCP probe (if
mcp__loom__*tools are registered in this session): callmcp__loom__list_sweeps. It returns a (possibly empty) registry on a healthy daemon, and normally fails fast if the IPC socket is missing or the process is dead. If nomcp__loom__*tools are registered at all, or the call takes more than a few seconds without returning (the historical #4043 wedge — unary MCP calls hanging up to 1800s), do not wait it out — go straight to step 2. - CLI probe (fallback, and equally valid as a first choice): run
loom-daemon status. It connects to the same running daemon over its Unix socket, with no MCP bridge in the path, and prints in-flight sweeps, the three dynamic-cap inputs, the main-health-gate halt state, and per-token usage — a superset of whatlist_sweepsreturns.
Call: mcp__loom__list_sweeps
# CLI fallback / equally-valid first choice:
loom-daemon status # human-readable table
loom-daemon status --json # machine-readable
loom-daemon status --pipeline # + forge-side pipeline snapshot (extra gh calls per managed repo)
If both probes fail (daemon unreachable)
Display this message and EXIT:
The Loom daemon is not running (both mcp__loom__list_sweeps and
`loom-daemon status` failed to reach it).
The daemon is a long-lived Rust process. Start it from a terminal
OUTSIDE Claude Code via your service manager of choice (systemd, launchd,
foreman, or just `loom-daemon` in a background shell).
While the daemon is down, in-process orchestration still works:
/loom:sweep <issue> # Single-issue lifecycle, in-session
# (subagent dispatch, single OAuth token)
Stage -1 of /loom:sweep auto-detects the daemon — when the daemon comes
back up AND a multi-account token pool is configured (.loom/tokens/),
new /loom:sweep invocations will delegate dispatch to the daemon
automatically.
If either probe succeeds (daemon reachable)
Proceed to the Observer / Dispatch Loop below. If only the CLI probe succeeded (no MCP tools available this session), use the CLI equivalents in that section throughout — the loop's underlying logic (assess pipeline, dispatch, monitor, cancel) is unchanged, only the tool surface differs.
CLI Fallback (Non-MCP Operator Surface)
The loom-daemon binary talks to the same running daemon over the same Unix-socket IPC that MCP tools use — it is not a separate, lesser code path, and it works in any shell regardless of whether mcp__loom__* tools are registered in the current Claude Code session. Verified against loom-daemon --help (each subcommand's --help) as of this writing; re-verify if the CLI surface changes.
| MCP tool | CLI equivalent | Notes |
|---|---|---|
mcp__loom__list_sweeps | loom-daemon status [--json] [--pipeline] | Richer than list_sweeps: also reports the three dynamic-cap inputs (disk headroom, ram headroom, configured ceiling) + their min, the main-health-gate halt state, and per-token usage. --pipeline adds the forge-side gh snapshot (opt-in — extra API calls). |
mcp__loom__dispatch_sweep | loom-daemon dispatch <N> [--model] [--effort] [--depends-on] [--workspace] | First-class non-MCP entry point (#3952) over the same IPC DispatchSweep request. Bounded client-side ack timeout — exits nonzero fast instead of hanging (built explicitly to avoid the #4043 MCP wedge). |
mcp__loom__cancel_sweep | loom-daemon cancel <sweep-id> | --issue <N> [--grace] [--workspace] | First-class non-MCP entry point (#4980) over the same IPC CancelSweep request — the dispatch sibling, usable over ssh. Do NOT kill -TERM <pid> from loom-daemon status instead (the pre-#4980 fallback): the daemon tracks the wrapper pid, so killing it leaves the underlying claude agent alive — on 2026-08-03 that survivor relaunched its workload against an issue whose claim had already been returned to the queue. The CLI signals the whole process group. .loom/sweep-checkpoint/ still survives a cancel, so redispatch still resumes. |
mcp__loom__get_sweep_status | (none, partial) | loom-daemon status gives fleet-wide state, not one sweep's phase/blockers. loom-daemon watch add <N> (add --pr to watch a PR instead of an issue) registers a durable watch on that issue/PR's terminal state instead (persists to ~/.loom/watches.json, survives a daemon restart, resolves to ~/.loom/logs/watch-results.log). |
mcp__loom__tail_sweep_log | (none) | tail -f .loom/logs/sweep-issue-<N>.log directly. |
mcp__loom__tail_event_bus, subscribe_to_events, publish_event | (none) | Live pub/sub is MCP-only; there is no CLI event stream. Poll loom-daemon status + gh issue/pr list on your loop cadence instead. |
CLI-only operator surface (no MCP equivalent at all — these predate or fall outside the MCP tool set):
loom-daemon restart[--drain] [--timeout] [--force-after-timeout] [--abort-drain] — deliberate supervised restart;--drainfinishes in-flight sweeps first (#4090)loom-daemon quarantine clear <issue>— release an insta-crash pause and restoreloom:issueon the forgeloom-daemon tokens {select,bootstrap,import-from-monitor,check,pin,unpin,unblock}— multi-account OAuth pool management (.loom/tokens/)loom-daemon watch {add,list,remove}— durable operator watches on issue/PR terminal stateloom-daemon workspace ...— the machine-level workspace registry (~/.loom/workspaces.json)loom-daemon stats— agent effectiveness/activity metrics
Note: the older loom-daemon --status / --health flag spellings are gone — these are subcommands now (loom-daemon status, no CLI --health at all; use loom-daemon status for daemon health, it folds in the main-health-gate state).
Observer / Dispatch Loop
When the daemon is running, you coordinate work via MCP tools where available, and the loom-daemon CLI everywhere else (or as a preference — the CLI is not degraded, just narrower in scope: it has no event-subscription/log-tail surface, since those are inherently a live-stream concept that fits an MCP session, not a one-shot CLI invocation).
CLI fallback quick reference (see "CLI Fallback" below for the full table with rationale):
| MCP tool | CLI equivalent |
|---|---|
mcp__loom__list_sweeps | loom-daemon status |
mcp__loom__dispatch_sweep | loom-daemon dispatch <N> |
mcp__loom__cancel_sweep | loom-daemon cancel <sweep-id> / --issue <N> |
mcp__loom__get_sweep_status, tail_sweep_log, tail_event_bus, subscribe_to_events, publish_event | none — MCP-only |
Each iteration:
-
Read current state:
- MCP:
mcp__loom__list_sweeps(currently-dispatched sweeps with PIDs and started_at),mcp__loom__get_sweep_status <sweep_id>(per-sweep phase, blockers, last activity),mcp__loom__tail_event_bus(short tail, recent lifecycle events) - CLI:
loom-daemon status(registry + dynamic-cap inputs + health-gate state + per-token usage; add--pipelinefor the forge-side snapshot in the same call, replacing step 2's separateghcalls)
- MCP:
-
Assess pipeline using read-only gh commands (or
loom-daemon status --pipelinefrom step 1):gh issue list --label="loom:issue" --state=open --json number,title --limit=20 gh issue list --label="loom:building" --state=open --json number,title --limit=20 gh pr list --label="loom:review-requested" --json number,title --limit=20 -
Dispatch new sweeps via MCP or CLI. Derive the target workspace root once and pass it explicitly — omitting it routes through registry resolution (#4299/PR #4322), which can silently target the daemon's default workspace instead of the repo you meant (#4503):
WORKSPACE_ROOT=$(git rev-parse --show-toplevel) For each ready loom:issue not already in the daemon registry: mcp__loom__dispatch_sweep kind={"Issue": <N>} workspace_root=$WORKSPACE_ROOT# CLI equivalent (#3952) — same underlying IPC DispatchSweep request: loom-daemon dispatch <N> --workspace "$WORKSPACE_ROOT" [--model M] [--effort E] [--depends-on P]kind/<N>is the only required input. Optional params on both surfaces:model,effort,depends_on(a single parent issue for stacked PRs); MCP additionally takesidempotency_key(dedup) and CLI additionally takes--workspace(target a non-default managed workspace root — always pass it explicitly rather than relying on the default). The daemon picks an OAuth token from the pool (spawn-claude.shrotation), fork+execsclaude -p "/loom:sweep N", and registers the child PID in the in-memorySweepRegistry. Token rotation only happens at this process-spawn boundary.loom-daemon dispatchapplies a bounded client-side ack timeout — if the daemon doesn't respond within a few seconds it exits nonzero with a clear error instead of hanging (the #4043 MCP-wedge failure mode this CLI path was built to avoid). -
Monitor lifecycle events (MCP-only, optional, for live debugging or stuck-sweep detection):
mcp__loom__subscribe_to_events --topic "sweep.issue.*"The frozen v0.10.0 topic taxonomy is:
sweep.issue.{N}.phase— phase transitions (curator → builder → judge → doctor → merge)sweep.issue.{N}.blocker— a sweep added aloom:blockedorloom:operator-onlylabelsweep.issue.{N}.exited— clean exit (withexit_codeandduration_sec)sweep.issue.{N}.crashed— non-zero exit / OOM (withcheckpoint_phase)sweep.global.dispatch— daemon accepted a newdispatch_sweeprequestsweep.global.completed— sweep completed (terminal state, post-reaper)
No CLI equivalent — if MCP is unavailable, poll
loom-daemon statusandgh issue/pr liston your ~30s loop cadence instead of subscribing. -
Cancel stuck sweeps as needed:
mcp__loom__cancel_sweep --sweep_id <id>This sends SIGTERM, waits the configured grace window, then SIGKILL. The
.loom/sweep-checkpoint/issue-<N>.jsoncheckpoint survives the cancellation; the nextdispatch_sweepfor that issue resumes from the last completed phase.CLI equivalent (#4980), for when MCP is unavailable or you are on ssh:
loom-daemon cancel <sweep-id> # or: loom-daemon cancel --issue <N>Same IPC request, same daemon-side termination. Do not
kill -TERM <pid>the PID fromloom-daemon statusinstead — that was the pre-#4980 fallback and it is how the 2026-08-03 incident happened: the tracked PID is the wrapper, so killing it leaves theclaudeagent alive to relaunch its workload against an issue whose claim has already been returned to the queue. If a sweep keeps crash-looping on redispatch,loom-daemon quarantine clear <issue>clears the daemon's insta-crash pause (crash-loop protection, not a cancellation substitute). -
Tail per-sweep logs if you need to inspect output:
mcp__loom__tail_sweep_log --issue <N> --lines 200Or use the bare-event-bus view:
mcp__loom__tail_event_bus --lines 50No CLI equivalent — tail the log file directly instead:
tail -f .loom/logs/sweep-issue-<N>.log -
Sleep ~30 seconds, then repeat.
Orchestration Logic
Normal autonomous operation:
- Count
loom:issueitems in the forge - Check active sweeps via
mcp__loom__list_sweeps - If issues are available and the daemon is not at capacity (operator-defined; the daemon itself does not enforce a hard limit), dispatch new sweeps
- If pipeline is empty (no issues, no proposals), prompt the operator to consider triggering Architect/Hermit manually — work-generation cadence is tracked under #3381 and is not dispatched by the daemon
- Monitor
sweep.issue.*.blockerevents for sweeps that added a blocker label; surface these to the operator - Monitor
sweep.issue.*.crashedevents for non-zero exits; consider re-dispatch (the checkpoint preserves progress)
Every dispatched /loom:sweep runs the full lifecycle and merges each
approved PR on Judge approval — there is no separate merge-gated mode and
mcp__loom__dispatch_sweep has no --force/--merge parameter.
Multi-account scaling
The daemon is the only path that gives autonomous orchestration multi-account OAuth token rotation:
- Each
mcp__loom__dispatch_sweepcall fork+execs a freshclaude -p "/loom:sweep N"child spawn-claude.shselects a token from.loom/tokens/.ranking(or the allowlist, or random fallback) and exportsCLAUDE_CODE_OAUTH_TOKENbefore exec- Multiple sweeps can run concurrently under different tokens, spreading load across accounts
In-session subagent dispatch (/loom:sweep with Stage -1 falling through to subagent path) inherits the parent's single OAuth token — fine for short batches, fatal for multi-day runs. The daemon path exists precisely to break that limit.
Commands Quick Reference
| Command | Description |
|---|---|
/loom:loom | Check daemon (MCP list_sweeps, fallback loom-daemon status), start observing/dispatching |
/loom:loom status | mcp__loom__list_sweeps, or loom-daemon status if MCP is unavailable |
/loom:loom health | Display daemon health summary (registry + recent events, or loom-daemon status which folds in the health-gate state) |
/loom:loom stop | Cancel all in-flight sweeps via mcp__loom__cancel_sweep (CLI: loom-daemon cancel <sweep-id>, #4980); daemon process itself stays alive |
/loom:loom help | Show comprehensive help guide |
/loom:loom help <topic> | Show help for a specific topic |
Cancelling sweeps and stopping the daemon
Cancel individual sweeps (preferred):
mcp__loom__cancel_sweep --sweep_id <id>
Cancel all in-flight sweeps:
For each sweep returned by mcp__loom__list_sweeps:
mcp__loom__cancel_sweep --sweep_id <sweep_id>
CLI equivalent (#4980), for a shell / ssh session with no MCP server:
loom-daemon cancel <sweep-id> # or: loom-daemon cancel --issue <N>
Same IPC request and same daemon-side termination as the MCP tool, and it signals
the whole process group. Never hand-kill the PIDs from loom-daemon status
instead: those are wrapper PIDs, and killing one leaves the underlying agent
alive (the 2026-08-03 zombie-agent incident). .loom/sweep-checkpoint/ files
still survive a cancel, so a later dispatch_sweep / loom-daemon dispatch for
that issue resumes from the last completed phase.
Stop the daemon process itself is out of scope for this skill — the daemon is a long-lived service that the operator manages outside Claude Code (via their init system, foreman, or shell-level process management).
HELP REFERENCE
When the user runs /loom:loom help, display the content below formatted as markdown. If the user provides a sub-topic (e.g., /loom:loom help roles), display only the matching section. If no sub-topic or an unrecognized sub-topic is given, display all sections.
Available sub-topics
List these when showing the full help or when the sub-topic is unrecognized:
/loom:loom help - Show this full help guide
/loom:loom help quick-start - Getting started in 60 seconds
/loom:loom help roles - All available agent roles
/loom:loom help commands - Slash command reference
/loom:loom help workflow - Label-based workflow overview
/loom:loom help daemon - Daemon mode and MCP-tool reference
/loom:loom help sweep - Single-issue orchestration
/loom:loom help worktrees - Git worktree workflow
/loom:loom help labels - Label state machine reference
/loom:loom help troubleshoot - Common issues and fixes
Sub-topic: quick-start
Getting Started with Loom
Loom orchestrates AI-powered development using GitHub issues, labels, and git worktrees.
Try it now - Manual Mode (one terminal per role):
# 1. Start as a Builder and work on an issue
/builder
# 2. In another terminal, review PRs as a Judge
/judge
# 3. Or curate issues to add implementation guidance
/curator
Try it now - Single Issue (sweep handles the full lifecycle):
# Orchestrate one issue from curation through merge
/loom:sweep 123
Try it now - Daemon Mode (multi-account autonomous dispatch):
# Step 1: Ensure loom-daemon is running (outside Claude Code, via your
# service manager). Verify via:
# mcp__loom__list_sweeps
#
# Step 2: In Claude Code, observe and dispatch:
# /loom:loom
#
# /loom:loom uses MCP tools to enumerate the registry, dispatch new sweeps,
# subscribe to lifecycle events, and cancel stuck work.
Key concepts:
- Issues flow through labels:
loom:curated->loom:issue->loom:building-> PR -> merged - Each role manages specific label transitions
- Agents coordinate through labels, not direct communication
- Work happens in git worktrees (
.loom/worktrees/issue-N) - Multi-account token rotation only works at process-spawn boundaries — that is the architectural reason daemon mode exists alongside in-session subagent dispatch
Sub-topic: roles
Agent Roles
Loom has three layers of roles:
Layer 2 - System Orchestration:
| Command | Role | What it does |
|---|---|---|
/loom:loom | Daemon | Observes the loom-daemon registry via MCP tools, dispatches sweeps via mcp__loom__dispatch_sweep, and monitors lifecycle events via the pub/sub bus. |
Layer 1 - Issue Orchestration:
| Command | Role | What it does |
|---|---|---|
/loom:sweep <N> | Sweep | Orchestrates a single issue through its full lifecycle: Curator -> Builder -> Judge -> Doctor -> Merge. Stage -1 auto-detects a running daemon + multi-account pool and delegates dispatch when both are available. |
Layer 0 - Task Execution (Worker Roles):
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 63
- Forks
- 9
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
loom-loom- Source
- github.com/rjwalters/kicad-tools