Contributing to CAO
SkillAI & modelsLets your agent contribute code changes to the CAO (CLI Agent Orchestrator) project.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Contributing to CAO skill
About this skill
Contribute changes to the CAO (CLI Agent Orchestrator) codebase, the local
What this skill tells your AI
The instructions your AI receives, as published by awslabs/cli-agent-orchestrator in skills/cao-contributing/SKILL.md and read by ahel’s review.
How to make a change to the cli-agent-orchestrator codebase and get it through CI
cleanly. Read this before pushing a branch or opening a PR. The canonical human docs are
DEVELOPMENT.md, CONTRIBUTING.md, and
CODEBASE.md — this skill is the operational checklist that mirrors what
CI actually enforces.
Golden rules (read these first)
- Run Python tooling through
uv.uv sync --all-extras --dev,uv run pytest …,uv run mypy src/,uv run cao …. There is no barepipworkflow, and the venv CI builds is the oneuvmanages. Repo scripts are documented in their own text as plainpython scripts/<name>.py(for examplescripts/sync_skills.py, whose fix-up message andtest_skill_packaging_parity.pyboth quote that form) — run them asuv run python scripts/<name>.py, which satisfies both. - Verify the actual CI run after every push — never declare "done" on local tests
alone. Poll it:
gh pr checks <number>for every workflow on the PR, orgh run list --branch <branch> --workflow CIforci.ymlalone, thengh run view <id>/gh run view <id> --log-failed. - When a required check fails unexpectedly, diff EVERYTHING your commit changed —
including CI/workflow/config files (
.github/workflows/*.yml,pyproject.toml,mypy.ini) — before concluding the cause is pre-existing or external. The signal is often in your own diff (git diff <base>..HEAD -- .github/). A displaced one-line workflow key (see the mypy note below) can turn a tolerated warning into a hard failure. - Never mark a task complete while a required CI gate is red. A red gate means not done; investigate, don't rationalize.
- Match the repo, don't reshape it. Don't bundle unrelated fixes (e.g. repo-wide type errors) into a feature PR, and don't tighten a CI policy as a side effect of an unrelated change.
Local dev loop
uv sync --all-extras --dev # install (mirrors what CI does)
uv run pytest test/path/to/test_x.py # run targeted tests while iterating
uv run black src/ test/ # format (CI checks --check)
uv run isort src/ test/ # import order (CI checks --check-only)
uv run mypy src/ # type check (see the mypy note below)
Write tests RED-first: add a test that reproduces the bug/behavior and fails, then implement until it passes. New features and bug fixes ship with tests.
Keeping patch coverage at 100% for changed lines is a team convention, not a gate.
Nothing fails your build over it: the repo has no codecov.yml, so there is no configured
status check or target, and the Unit Tests job uploads coverage with
fail_ci_if_error: false — even a broken upload is tolerated. Treat the Codecov comment as
a review signal to justify, not a red gate to chase.
The CI gate map (.github/workflows/ci.yml)
Know which jobs are blocking vs tolerated so you can tell a real failure from
noise. test/test_cao_contributing_skill_accuracy.py fails if this table drifts from
ci.yml, so trust it — and if you rename a job, update it here.
| Job | Runs | Blocking? |
|---|---|---|
| Unit Tests (3.10 / 3.11 / 3.12) | uv run pytest test/ examples/workflow/tests/ --ignore=test/providers/test_kiro_cli_integration.py --ignore=test/e2e -m "not e2e" --cov=src/cli_agent_orchestrator --cov-report=term-missing | Yes |
| ↳ step: Validate Markdown links | uv run python scripts/validate_markdown_links.py — every relative link in every tracked .md, including skills/ | Yes |
| Code Quality | black --check, isort --check-only, then uv run mypy src/ | black/isort yes; mypy is non-blocking (continue-on-error: true) |
| AG-UI demo (shift-left recording) | boots a CAO_AGUI_ENABLED server, drives the viewer, records a GIF artifact | Yes |
| AG-UI construct demos (shift-left recordings) | same pattern for the L2 construct library | Yes |
| AG-UI stock-client live (AC3) | drives a real third-party AG-UI client against the surface | Yes |
| Agent Plugins dog-food (shift-left recording) | records the plugin pipeline from examples/agent-plugins/agent-plugins-dogfood/tools and gates on drift | Yes |
| CAO MCP Apps | MCP Apps build + backend coverage ratchet floor | Yes |
| CAO MCP Apps E2E (Playwright) | browser E2E over the ui://cao/* views | Yes |
| Rust TUI (Linux x86_64 / macOS arm64) | cargo test for the tui/ crate | Yes |
| Web UI Build | frontend build | Yes |
| AI-DLC Portfolio Example | example project builds | Yes |
| Security Scan | Trivy | Yes |
| Dependency Review | actions/dependency-review-action over the PR's dependency delta: fail-on-severity: high plus denied licences GPL-3.0/AGPL-3.0 | Yes — CI-only; there is nothing to run locally, and it is skipped on forks (if: github.repository == 'awslabs/cli-agent-orchestrator'), so a green run on your fork has not exercised it |
The
-m "not e2e"on the CI command replaces your localaddopts— it does not compose with it. So a local run that also deselectsintegrationis a strict subset of CI's selection and can be green while CI is red. Compare deselected counts, not just pass counts.CI's selection is narrower than the marker alone implies. The same command passes
--ignore=test/providers/test_kiro_cli_integration.pyand--ignore=test/e2e, and--ignorewins over-m: that Kiro provider test needs an authenticated external CLI and never runs in CI. So mostintegration-marked tests do run in CI — but not that one, and nothing in CI covers it. If you change it, you have to run it yourself.
mypy is intentionally non-blocking. The repo has known, pre-existing, repo-wide mypy errors (historically in
services/agent_scaffold.py,cli/commands/profile.py,services/memory_service.py/api/main.pyMemoryArchiveBackendcall-arg, and ajsonschemastub). CI tolerates them viacontinue-on-error: trueon the mypy step. Therefore:
- Do not make mypy blocking, and do not bundle those unrelated type-fixes into a feature PR (they belong in a dedicated cleanup PR).
- Your change must add zero new mypy errors — check the delta, not the raw count.
- When you insert a new job/step near the
lintjob, keepcontinue-on-error: trueattached to the mypy step. A mis-insertion once displaced that line onto the next job's step, silently turning mypy into a hard gate and failing the build for unrelated, pre-existing errors.
Other workflows that gate a PR
ci.yml is not the only workflow on a pull request. These run alongside it, and none sets
job-level continue-on-error, so every check they run is blocking. The same accuracy test
pins this table to the workflow files.
| Workflow | Checks on the PR | Runs on | Blocking? |
|---|---|---|---|
Secret Scan (secret-scan.yml) | gitleaks, gitleaks config tests | every PR to main | Yes — the config tests run locally as uv run pytest test/test_gitleaks_config.py; the scan itself is gitleaks detect --config .gitleaks.toml over the PR's commits |
cargo-deny (cargo-deny.yml) | cargo-deny (advisories, licenses, bans, sources) | every PR to main | Yes — locally, cargo deny --manifest-path tui/Cargo.toml --locked check (global flags before the subcommand, as the action passes them) |
Test Antigravity CLI Provider (test-antigravity-cli-provider.yml) | Unit Tests, Code Quality | only PRs touching that provider, its unit test or fixtures, pyproject.toml, or the workflow | Yes |
Test Claude Code Provider (test-claude-code-provider.yml) | Unit Tests, Code Quality | only PRs touching that provider, its unit test, pyproject.toml, or the workflow | Yes |
Test Codex CLI Provider (test-codex-provider.yml) | Unit Tests, Code Quality | only PRs touching that provider, its unit test or fixtures, pyproject.toml, or the workflow | Yes |
Test Kiro CLI Provider (test-kiro-cli-provider.yml) | Unit Tests, Code Quality | only PRs touching that provider, its unit test or fixtures, pyproject.toml, or the workflow | Yes |
Docs site (gh-pages.yml) | build | only PRs touching docusaurus/** or the workflow | Yes — deploy is push-only and never runs on a PR |
Two checks can share a name. Each provider workflow has its own
Unit TestsandCode Quality, so a PR that touchespyproject.tomlshows those names more than once. Read the workflow name next to a failing check before assuming it is the CI one —gh pr checks <number>lists every workflow, whereasgh run list --workflow CIsees onlyci.yml.
Testing gotchas
- The full
uv run pytest test/is flaky locally — it needs a running server, tmux, and real CLI binaries, and can hit a flaky OTel/gRPC abort. Run targeted test files while iterating and lean on CI (the Unit Tests job) for breadth; get authoritative missing-coverage lines from that job'sterm-missingoutput ∩ your diff. CI is not the full suite, though — it excludestest/e2eand the Kiro provider integration test by path (see the callout above), so those two are only ever covered by someone running them deliberately. - A test that touches CAO's own state can pass only on your machine. If a code path
reaches the real database, config dir, or a live session, it will be green on a
developer box that has an initialised CAO install and red on a clean runner with
sqlite3.OperationalError: no such table: terminals. Mock the store/DB seam explicitly; when a test exercises a service function, check what that function calls today — a rebase can introduce a new unmocked DB write into a path your test already covered. Verify by running under a throwawayHOMEand a throwawayCAO_HOME_DIR—HOMEalone is not enough, becauseconstants.pyprefers an exportedCAO_HOME_DIRand derives the database and every other state path from it, so an absolute value you already export keeps the initialised store this is meant to exclude:
Keep the trap-in-a-subshell shape, and keep the diagnostic inside it. Cleaning up withTMPH=$(mktemp -d) ( trap 'rm -rf "$TMPH"' EXIT # cleans up on every path HOME="$TMPH" CAO_HOME_DIR="$TMPH/cao" \ uv run pytest test/path/to/test_x.py rc=$? # not `status`: read-only in zsh echo "exit=$rc" # pytest's status, not rm's exit "$rc" ) # ...and the block returns it too; rm -rf "$TMPH"returnsrm's status instead of pytest's, so a failing run reports success; switching to&&fixes the status but leaks the temp directory on exactly the failures you wanted isolated. A trailingecho "exit=$?"after the subshell prints the right number but is itself the block's last command, so the block returns 0 and a script or agent checking$?still reads a failing run as a pass. - Local green and CI green are different claims, in both directions. A local suite can hide real failures (see above) and invent ones CI never sees (macOS-only, missing optional binaries). When they disagree, CI is authoritative — read the job log rather than reasoning from the local result.
- FastAPI
TestClientmust usebase_url="http://localhost"— the Host-header / DNS-rebinding guard returns400otherwise. - Provider status detection is screen-scraping — provider tests are fixture-driven state machines; when a CLI tool changes its TUI, update the regexes and add a fixture.
- The AG-UI demo recorder (
examples/ag-ui/ag-ui-eventsource-viewer/tools) needs a Chromiumheadless_shellmatching the pinned@playwright/testversion (npm run playwright:install) plusffmpeg, and boots its ownCAO_AGUI_ENABLEDserver. It gates in CI, so you don't have to run it locally to land a change. The construct-demo recorder is the sibling atexamples/ag-ui/ag-ui-construct-demos/tools.
Pre-PR checklist
uv run black src/ test/ && uv run isort src/ test/(or--checkto verify).uv run mypy src/— confirm no new errors vs the base (pre-existing ones are OK).uv run pytest <targeted files>green; add/keep tests for changed behavior.- If you touched any
.md, runuv run python scripts/validate_markdown_links.py— a dead relative link fails the Unit Tests job, andskills/is in scope. CAO has no rootAGENTS.md; the contributor map isCODEBASE.md. - If you touched
skills/, runuv run python scripts/sync_skills.pyso the packaged mirror stays in lockstep (test/test_skill_packaging_parity.pyenforces it). - Commits: only when asked; sign if the repo expects it; keep the subject concise and
Conventional-Commits style; never force-push to
main. - Open the PR, then watch its CI run to completion and fix any red gate before calling
it done (rule #2 and #4). Use
gh pr create/gh pr checks.
Not what you want?
- Authoring a new agent skill (
SKILL.md, frontmatter, evals) → no shipped skill covers this yet; follow the Agent Skills specification directly. - Building a provider / plugin / MCP-apps view → cao-provider / cao-plugin / cao-mcp-apps.
- Launching or steering running agent sessions → cao-session-management.
Signals
- GitHub stars
- 1k
- Forks
- 277
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
cao-contributing- Source
- github.com/awslabs/cli-agent-orchestrator