Adding Tests and CI
SkillAI & modelsWhere a new test belongs and how to scope CI jobs, steps, and workflow triggers so they run only when a change can affect them. Use when adding or moving a test, adding or changing a CI job, step, or workflow, or when a job is slow or runs on unrelated changes.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Adding Tests and CI skill
What this skill tells your AI
The instructions your AI receives, as published by builderio/agent-native in .agents/skills/adding-tests-and-ci/SKILL.md and read by ahel’s review.
Why this is a cost question
Every agent-native workflow draws on one org-wide pool of GitHub-hosted runners.
On the Team plan it held 60 concurrent jobs, and on 2026-09-29 the Fast tests
gate waited a median 445 s just for a slot. BuilderIO moved to Enterprise
Cloud in October 2026, which raises the pool to 500 jobs (50 macOS). Minutes
are free for this OSS repo; slots are shared with every private repo. Scope
still matters, because a job a change cannot affect delays other PRs once the
pool fills, but parallelism that cuts the critical path is now worth a slot.
scripts/ci-change-scope.ts is the one place that decides what a change runs.
It classifies the changed paths into a full or targeted run and emits one output
per check. ci.yml jobs gate on those outputs. A .github/** change, the
lockfile, root config, or a scripts/** change other than a guard script or
root script test forces a full run.
Adding a test
- Put it in the workspace of the code it proves. Targeted runs test the
changed packages and their dependents (
...{packages/core}), so a test in another package either runs on unrelated changes or never runs on the one that breaks it. - Prove it at the cheapest tier that can fail for the bug: a pure
function, then a component or
.db.test.ts, then browser E2E last. Some E2E suites, such as Design's, run only after merge and cannot fail a PR. Do not restate one assertion at several tiers.design-editor-architecturecovers failing-first proof and mutation audits. - A new test joins an existing lane or job. It never gets a job of its
own. A changed root
scripts/*.test.ts, or the test beside a changed guard script, runs inSecurity guardsthrough thescript_testsoutput. - Mind what
vitest --changedcannot see. Targeted Core lanes run only the tests whose import graph reaches a changed file. A test that reads a file from disk (a skill, a fixture, a template) will not rerun when only that file changes. Import it, or teachrequiresFullCoreFastTestsinscripts/ci-test-lanes.tsabout the path. - Never skip, disable, or quarantine a test to get green.
Test workers
Fast-test lanes run Vitest with VITEST_CONCURRENCY=100%, one worker per
core. The shared config's 25% default is for laptops, and on CI's 4-vCPU
runners it meant one worker: a core shard took 542 s at one worker and 201 s
at four, on the same CPU time. On the 60-slot pool that let two full-width
lanes replace five, but each lane then ran ~18 min of packages one after
another. With 500 slots CI plans eight lanes (LANES in ci.yml): each costs
~100 s of setup, and the planner balances lanes by test-file count.
-
Large packages are sharded too.
splitLargePackagesinscripts/ci-test-lanes.tssplits a package heavier than a fair lane share into Vitest--shards, so Design no longer sets the floor for every lane. Only a barevitesttest script can be sharded, and no shard drops belowMIN_SHARD_FILES. -
Let CI override the worker count. A package that pins
maxWorkersgoes throughresolveMaxWorkers(process.env, fallback); a literal overrides the CI setting. -
Leave headroom for work inside a test body. Booting PGlite or importing a server bundle takes several times longer when every core is busy, and a different test crosses the limit each run. The shared config allows 30 s; do not add a stricter per-block timeout.
-
Give parallel workers separate databases. A file that opens the database without
DATABASE_URLgets the default directory, whose lock admits one process at a time. Core's Vitest setup gives each worker slot its own under a root that belongs to one run, so concurrent runs never share one; a test that asserts the default URL clearsDATABASE_URLitself. -
Keep the forks pool with isolation. Threads ran slower and failed 79 tests; turning isolation off gained 7% and failed 129.
Adding or changing a CI job
| Rule | Why, from this repo |
|---|---|
Give a job its own check in CHECK_NAMES whose condition names the paths that can change its result. Never borrow another job's output. | The Postgres connection budget reused neon_query_budget and ran on every template edit. Its probe imports only @agent-native/core/db. |
| Per-app work takes a list output, not a boolean. Shared packages or a full run select every app; a template change selects that template; an empty selection turns the check off. If the check is on and the list is empty, the job fails. | query_budget_apps and ssr_boot_apps: a one-template PR built and measured all 16 templates (12–13 min). When shared code selects every template, query_budget_matrix splits them across two jobs. The beta publisher's discover-sites publishes only sites whose dependency closure changed. |
| Split the expensive step from the cheap one. Keep the cheap check on every run and gate the expensive step on its narrower inputs. | Android: expo export runs whenever a file Metro bundles changes, but the Gradle compile, the bulk of the ~23 min Android job, runs only when packages/mobile-app, the lockfile, or the workflow changed. |
A path filter covers the job's whole dependency closure, including install-time inputs: root package.json, pnpm-workspace.yaml, the prebuild script, and the lockfile. | The desktop canary filter missed the bundled Chrome extension and then the postinstall inputs. Each gap skipped runs that should have caught a break. guard:mobile-build-paths traces Metro's imports and fails when the mobile filter misses one. |
Subscribe only to events that can change the outcome. Validators pin some trigger sets, e.g. scripts/validate-*-workflow.ts and scripts/package-release-workflow.test.ts; update them in the same change. | Content product conformance reran on labeled/unlabeled but never reads labels, and on title edits although it reads the body only for its declaration. |
Do not add a job for under a minute of work. Checkout, install, and a pool slot cost more than the check. Fold it into a job with the same setup, and do not duplicate what pnpm guards already runs. | Consolidation took PR pushes from ~31 to ~23 jobs. It folded PGlite locking into Content DB tests, privacy evals into Brain evals, QA static into Security guards, and the changeset check into Lint & format, and dropped a drizzle guard job pnpm guards covered. |
| Batch periodic publishing on a schedule with change detection, instead of once per merge. | Nightly npm snapshots publish every 3 h, and only when a publishable path changed since the last successful scheduled run. |
PR workflows cancel superseded runs: group: <name>-${{ github.event.pull_request.number || github.ref }} with cancel-in-progress: true. Publishers on main use cancel-in-progress: false so a deploy is never killed mid-flight. An event rejected only by a job if still joins the run's group and cancels a real run, so filter it in on: or give ignored events a throwaway group. | Visual Recap events the gate ignored used to cancel an in-progress recap; they now get a throwaway group. |
| When scope cannot be computed, run everything or fail. Never skip. | An unreadable beta site list, an unknown base, or unrelated history publishes every site. A failed git diff fails the step; it is never || true. |
| Print what fails, not everything. | The full-tree lint printed ~15k pre-existing warnings over the one error. It now runs oxlint --quiet. |
A skipped job counts as passing for required checks, so gating a required job
or step with if is safe. Every uppercase key a workflow adds under env:
needs a docs/environment-variables.md row (guard:env-documentation). Prefer
built-in GITHUB_* variables over new ones.
Proving a CI change
- Classifier: add
scripts/ci-change-scope.test.tscases for each new check: on, off, full run, and tooling-only full run. Then runnode --experimental-strip-types --test scripts/ci-change-scope.test.ts scripts/ci-build-workspaces.test.ts scripts/ci-test-lanes.test.ts. - Real history: run
scripts/ci-change-scope.tswithGITHUB_OUTPUTset on realmaincommit ranges, and put the before/after in the PR. - Shell: run each new
run:snippet locally against a matching input and an empty one. Runactionlinton the workflow when it is available. - Mind the self-test gap: a PR that edits
ci.ymlruns full CI itself, so its own checks cannot show the saving. The classifier output is the evidence.
Real failures this replaces
- "whats up with Cold-request query budget. runs on every CI for every template? seems wrong"
- "fast lane tests seem absuredly expensive"
- "why is Generate + run standalone Chat running so frequently? shouldnt it listen to Chat changes?"
- "so tldr: ~2-3 simultaneous PRs will eat up our entire quota of jobs right"
- "can you make lint & format not print all the warnings?"
Signals
- GitHub stars
- 7k
- Forks
- 635
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
adding-tests-and-ci- Source
- github.com/builderio/agent-native
github.com/builderio/agent-native
Related picks
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptanimate-expo
Skill · emilkowalski
The pick for Expoexpo-dev-client
Skill · expo
The pick for Exposupabase-postgres-best-practices
Skill · supabase
The pick for Postgreshandsontable-playwright-e2e
Skill · handsontable
The pick for End-to-end testing