agent-council

SkillAI & models

Use when facing trade-offs, subjective judgments, uncertain decisions, or when diverse viewpoints would improve judgment quality. Triggers include "council", "다른 의견", "perspectives", "what do others think".

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the agent-council skill

What this skill tells your AI

The instructions your AI receives, as published by toongri/oh-my-toong-playground in skills/agent-council/SKILL.md and read by ahel’s review.

Agent Council

Advisory body providing multiple AI perspectives on uncertain decisions. When all members are unavailable, falls back to in-session single-voice advisory using the in-session fallback framework.

Council provides opinions. The caller makes the final decision.

Quick Reference

SituationCouncil Needed?Reason
Architectural trade-offs✅ YesMultiple perspectives needed
Subjective code quality judgments✅ YesNo clear correct answer
Risk assessment disagreements✅ YesDepends on perspective
Compilation/syntax errors❌ NoAn objective solution exists
Code style❌ NoHandled by ktlint
Clear spec requirements❌ NoOnly implementation needed

When to Use vs When NOT to Use

digraph council_decision {
    rankdir=TB;
    node [shape=box, style=rounded];

    start [label="Decision needed?", shape=diamond];
    objective [label="Objective answer exists?", shape=diamond];
    skip [label="Skip council\nDecide directly"];
    use [label="Use council"];

    start -> objective [label="yes"];
    start -> skip [label="no decision"];
    objective -> skip [label="yes (compile error, lint)"];
    objective -> use [label="no (trade-offs, judgment)"];
}

Use Council:

  • Architectural trade-offs (monolith vs microservice, sync vs async)
  • Subjective code quality decisions
  • Multiple valid approaches exist
  • Risk assessment disagreements

Skip Council:

  • Compilation/syntax errors (objective fix)
  • Code style (ktlint handles)
  • Clear spec requirements (just implement)
  • Time-critical simple fixes

Process

  1. Encounter uncertain decision point
  2. Call council with rich context + specific question
  3. Council members provide independent opinions (raw outputs)
  4. Collect: bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" collect JOB_DIR — polls internally; re-call if not "done"
  5. Read each member's output file via the Read tool 5a. Completeness gate (Chairman judgment): done does NOT mean semantically complete. A member can exit cleanly yet return a non-answer — a plan, framing, or "I'll answer once X arrives" response. Read each member's content and judge: did it actually answer the asked question? Also check: is any member in awaiting_resume state? For any member that is awaiting_resume OR whose content is a non-answer (narrative-only / incomplete / waiting pattern), call resume-member BEFORE proceeding. clean is destructive (it deletes the jobDir, and resume-member <jobDir> requires that jobDir) — so clean is ALWAYS the last step, only after completeness is confirmed.
  6. Synthesize (you as Chairman): raw outputs → Advisory Format below
  7. Make informed decision based on advisory
  8. Cleanup: bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" clean JOB_DIR

Context Synchronization

Council members do not share the caller's session context. The caller must explicitly provide:

  • Evaluation Criteria: Key principles from the review/validation rules
  • Project Context: Conventions and patterns discovered during session
  • Target Content: Code, spec, or artifact under review
  • Specific Question: Points where judgment is needed

Include context richly. Council members should judge with the same context as the caller.

How to Call

Execute bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" from the project root:

Note: Always write the council prompt in English for consistent cross-model communication.

Host Agent Context (Claude Code)

For programmatic use within Claude Code sessions:

CRITICAL: Always set timeout: 180000 on every Bash tool call.

Hard Constraints: 0. Each Bash call MUST run in FOREGROUND. All subcommands (start, collect) run synchronously. No background execution. No run_in_background.

  1. Do NOT use sleep. The collect subcommand polls internally (5-second intervals, 20-second wait per call). External sleep is redundant and wastes time.
  2. Exactly-once job start. The start subcommand runs ONCE. Polling (collect) may repeat. No job re-creation.

1. Start council (Bash, timeout: 180000)

PROMPT_FILE=$(mktemp)
cat > "$PROMPT_FILE" << 'PROMPT_EOF'
## Evaluation Criteria
[Key principles - in English]

## Project Context
[Conventions and patterns - in English]

## Target
[Code or content under review]

## Question
[Specific points needing judgment - in English]
PROMPT_EOF
JOB_DIR=$(bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" start --stdin < "$PROMPT_FILE")

Output: JOB_DIR path (one line on stdout).

Important: Write prompts in English for consistent cross-model communication.

2. Collect results (Bash)

One collect call is one poll: it polls internally every 5 seconds and answers within 20 seconds either way. Members take minutes, so repeated polls are the normal path. No external sleep needed.

bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" collect "$JOB_DIR"
  • If response shows "overallState": "done" → proceed to Step 3.
  • Otherwise ("running", "queued", etc.) → call collect again (same command, foreground).

3. Read raw outputs

Use the Read tool to read each member's outputFilePath from the manifest. Only read entries where outputFilePath is non-null (null = infrastructure failure; see Degradation Policy).

4. Completeness gate — resume-member if needed (Chairman judgment)

done does NOT mean semantically complete. Read each member's content. If any member is in awaiting_resume state OR returned a non-answer (narrative-only / incomplete / waiting pattern / "I'll answer once X arrives"), call resume-member before proceeding:

bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" resume-member "$JOB_DIR" <name> "Please answer the question directly."

The prompt is written by the Chairman LLM for the specific situation. The example above is illustrative only.

Cap: max 3 resumes per member. After cap exhaustion — partial-accept (include what was received) OR escalate the entire job. The Chairman judges which applies based on the situation.

WARNING: clean is destructive — it deletes the jobDir permanently, and resume-member <jobDir> ... requires that jobDir to exist. clean is ALWAYS the last step, called only after completeness is confirmed for all members.

5. Synthesize (caller responsibility)

You as the Chairman must synthesize raw outputs into the Advisory Format (see below). The council does NOT produce a synthesized advisory automatically.

6. Cleanup (Bash, timeout: 180000)

bun "${CLAUDE_SKILL_DIR}/scripts/job.ts" clean "$JOB_DIR"

CLI Flags — Chairman Inclusion

  • --include-chairman / --include-chairman=true: Force-include the chairman as a member regardless of config.
  • --exclude-chairman / --exclude-chairman=true: Force-exclude the chairman from members.
  • --exclude-chairman=false: Explicitly request the chairman BE included (value-respecting).
  • Without any flag: falls back to exclude_chairman_from_members in config. The semantic default is true (chairman excluded) when the config key is absent; shipped configs override this to false to include the chairman as a member.

Flag values are respected (value-respecting parsing). --exclude-chairman=false does NOT exclude the chairman — it explicitly keeps them.

When both --include-chairman and --exclude-chairman are passed simultaneously, --exclude-chairman wins (override precedence).

Synthesis Protocol

When synthesizing raw outputs:

  1. Extract each reviewer's core position and key reasoning
  2. Overlapping positions → Consensus
  3. Conflicting positions → Divergence (report ALL, not majority)
  4. Unique concerns from any reviewer → include in advisory
  5. On divergence: consult Model Characteristics table to inform weighting. Cite the table when weighting one model's opinion higher.

Model Characteristics (Synthesis Weighting)

Last verified: 2026-02 (review quarterly as models update)

When synthesizing, weight each model's opinion based on the question domain:

MemberPrimary StrengthsWeight Higher When
claudeNuanced trade-off reasoning, instruction coherence across long context, risk/impact assessmentArchitecture decisions, requirement ambiguity resolution, risk evaluation
codexCode-level feasibility analysis, implementation cost/complexity estimation, API contract design"Is this buildable?" questions, implementation approach choices, technical debt evaluation
glmIndependent reasoning from a different model lineage, alternative solution discovery, assumption challengesTechnology comparisons, "what are we missing?" questions, cross-checking consensus

Application rules:

  • On consensus: model strengths are irrelevant — report agreement as-is
  • On divergence: reference the table above. If the question is about implementation feasibility and codex disagrees with claude and glm, state: "Codex's position carries additional weight here as an implementation feasibility question (see Model Characteristics)"
  • On contradiction with table: if a model gives a strong argument outside its listed strengths, the argument's quality overrides the table. Strengths are tie-breakers, not vetoes
  • Never discard a model's opinion solely because the domain doesn't match its listed strengths

<Output_Format>

Advisory Output Format

Chairman synthesizes council opinions into:

## Council Advisory

### Consensus

[Points where council members agree]

### Divergence

[Points where opinions differ + summary of each position]

### Recommendation

[Synthesized advice. When model opinions diverge, note which model's expertise is most relevant to this domain and why — referencing Model Characteristics table]

</Output_Format>

Result Utilization

Strong Consensus → Adopt recommendation with confidence

Clear Divergence → Options:

  • Flag as "Clarification Needed"
  • Choose majority position, noting dissent
  • Use divergence to identify edge cases

Mixed Signals → Weigh perspectives based on relevance

Partial Results → Apply Degradation Policy. Synthesize from available responses, note missing perspectives.


Common Mistakes

MistakeWhy It's WrongFix
Calling council for compilation errorsAn objective solution exists; wastes timeFix directly
Sending only a question without contextJudgment is impossible without contextInclude evaluation criteria and project context
Accepting the council's decision as-isCouncil advises; the caller decidesConsider the opinions and make your own decision
Calling council for every decisionUnnecessary overheadUse only for trade-offs/subjective judgments
Calling council in KoreanReduces consistency across modelsWrite prompts in English

Red Flags - STOP Before Calling Council

Red FlagReality
"빨리 결정해야 해서 council 생략"More important decisions need more diverse perspectives
"내 판단이 맞으니까 확인만"Confirmation bias — council is used to hear counterarguments
"에러 메시지가 뭔지 모르겠어서"Objective problems are outside council's scope
"council이 결정해줄 거야"Council advises; the caller is responsible for the decision

Degradation Policy

Council members may fail due to CLI unavailability, timeout, or errors. This is NOT the same as quorum logic.

Critical distinction:

  • PROHIBITED quorum logic: "2/3 responded, that's enough, skip the third" — this is giving up on a working member
  • PERMITTED degradation: "2/3 responded, third member's CLI crashed (missing_cli/timed_out/error state)" — this is handling infrastructure failure

Decision tree:

  1. overallState === 'done' AND all members have terminal states?
  2. Check each member's state and output content:
    • missing_cli → CLI not installed. Degradation applies (infra failure — unrecoverable).
    • timed_out → CLI exceeded timeout. Degradation applies (infra failure — unrecoverable).
    • error (non-zero exit) → CLI failed. Degradation applies (infra failure — unrecoverable).
    • canceled → Manually stopped. Degradation applies (infra failure — unrecoverable).
    • awaiting_resume → Member paused mid-turn waiting for continuation. Recoverable — call resume-member (see completeness gate above). This is distinct from infra failure; the member is alive and can continue.
    • done but content is a non-answer → Member exited cleanly but returned framing / a plan / a waiting pattern instead of actually answering the question. Recoverable — call resume-member. The runner marks it done based on exit reason, not semantic completeness; the Chairman judges semantic completeness by reading the content.
  3. After all recoverable states are resolved (or resume cap exhausted), synthesize from members with substantive answers.

If start exits non-zero or returns an empty $JOB_DIR, skip Steps 2–4 (no collect/clean) and fall back to in-session advisory: READ prompts/default.md and apply the framework defined there IN-SESSION. You become the in-session advisor for the remainder of this skill invocation — no manifest exists to read. When the cause is the no-members guard — stderr contains to dispatch — enter the fallback silently (expected path: empty/all-filtered members config). For any other non-zero exit (an unexpected failure — disk/permission, spawn error, a bug), first surface the failure reason (include the stderr line) in your output, then proceed with the in-session fallback.

Synthesis by response count:

ResponsesActionOutput Modification
3/3Full synthesisStandard advisory format
2/3Partial synthesisPrepend: "⚠️ Partial advisory (2/3 respondents). [failed_member] unavailable: [state]. The following synthesis lacks [failed_member]'s perspective ([see Model Characteristics for what this model typically contributes])."
1/3Single response reportPrepend: "⚠️ Limited advisory (1/3 respondents). [failed_members] unavailable. Presenting single response from [available_member] without synthesis. Treat as individual opinion, not council advisory."
0/3In-session fallbackREAD prompts/default.md and deliver in-session advisory. When all members failed due to infrastructure errors (not the no-members guard), surface a one-line failure summary before proceeding.

Partial synthesis rules:

  • Use "partial consensus (N/3 respondents)" when reporting agreement
  • In Divergence section, note: "Note: [missing_member]'s perspective is absent. Based on Model Characteristics, this model typically contributes [strength area] — this gap may affect the advisory's completeness in that domain."
  • Do NOT extrapolate what the missing model "would have said"

Signals

GitHub stars
25
Forks
1
Last commit
Sep 2026

ahel review

  • K6low
    bundled executables the agent is told to run

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
agent-council-toongri
Source
github.com/toongri/oh-my-toong-playground