Maintain

SkillDev tools

Maintain the slopus/happy open source project. Triage issues, manage the GitHub project board, draft closing comments, find duplicates, check if bugs are fixed on main, and engage with community contributors. NEVER posts comments or closes issues without showing exact text and getting approval first

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Maintain skill

What this skill tells your AI

The instructions your AI receives, as published by altman-conquer/agentrejoin in .agents/skills/maintain/SKILL.md and read by ahel’s review.

Keep the dbt project correct as the world underneath it moves. Maintenance is the recurring half of the loop: warehouses drift, loads half-fail, models go stale, keys stop being unique, and business definitions change. This skill compares a known-good baseline against current reality, classifies what drifted, and proposes the reconciling edit. It is manual and on-demand here; continuous drift detection and automated PRs are the commercial product.

The model: baseline, detect, reconcile

Drift is measured against a baseline (the .dex/snapshot.json fingerprint of the warehouse map and the project's per-layer definitions). Detection is read-only; only reconcile proposes edits.

Snapshot discipline matters. A snapshot is only as trustworthy as the moment it froze. Take one right after a known-good build (maintain snapshot), and commit .dex/snapshot.json like a lockfile so the whole team diffs against the same reference. Snapshot a state that is already drifted and check will mask the very drift you care about. When you accept a change as the new normal (re-run explore map first, then maintain snapshot); check warns when the baseline looks stale.

On a warehouse past the rank cutoff, use explore map --full before snapshotting. Past 50 objects explore map profiles the top 25 by rank and enters the rest as metadata alone, and the baseline can only compare columns for objects it has columns for. Snapshotting a partial map is still valid, and the envelope reports column_detail_count against dataset_count plus a warning naming what it could not cover, so the gap is visible rather than silently mistaken for a clean bill.

How to drive it

uv run --no-project --script "${CLAUDE_SKILL_DIR}/scripts/run.py" <subcommand> [flags]

dex runs its engine through uv, which is a prerequisite and is not installed by Claude Code. If the shell reports uv: command not found, stop and tell the user to install it (curl -LsSf https://astral.sh/uv/install.sh | sh, or brew install uv, or pipx install uv), then re-run. Never fall back to diffing the warehouse against the project by hand instead: the drift axes and the baseline comparison live in the engine, so any other path is guesswork.

The first command in a fresh environment installs the engine, so it can take tens of seconds where later ones take well under a second. --warm pays that install up front and exits without running anything:

uv run --no-project --script "${CLAUDE_SKILL_DIR}/scripts/run.py" --warm

Offer it once at setup. It is not something to run before an ordinary command.

  • maintain snapshot captures or refreshes the baseline. Run it after a clean explore or transform session so later runs have a known-good reference. It pins the current .dex/cache.json (so the grain baseline is the exact-distinct verdicts explore map already computed) plus per-layer fingerprints of the dbt project. Without a cache it captures a metadata-only baseline and says so. It also warns when the cache it pinned is thin (objects without column detail) or older than the profile freshness window, because either makes an "accept current state" only partly true.

  • maintain snapshot --project-only is for a project-only refactor, such as moved model files or a dbt project rename. It refreshes transform and semantic fingerprints without opening the warehouse, carrying the previous warehouse evidence and original capture time forward instead. It requires an existing snapshot and refuses connection-target flags.

  • maintain check is the everyday entry point: it sweeps every axis and returns a report ranked by blast radius. Read-only.

  • maintain schema [<objects>] detects structural drift: source columns and tables added, dropped, retyped, or renamed; nullability changes; declared sources the warehouse no longer honors.

  • maintain volume [<objects>] detects freshness drift: row counts that collapsed, spiked, or went to zero. This is the "is the data still flowing correctly?" axis, distinct from "did the shape change?".

  • maintain grain [<objects>] detects grain drift: a key that now has duplicates, a changed row-per-entity cardinality, or an increased join fanout. It also re-verifies the grains your project declares (a model-level unique_combination_of_columns), which measurement on its own can miss. Uses aggregates, never raw rows.

    Two findings come out of the uniqueness checks and the difference is the baseline. key_lost_uniqueness is a key that was proven unique and is not any more: something changed in the data. declared_grain_not_unique is a declared combination that does not hold, and nothing changed at all: the project asserts a grain the data never had, so the fix is to the declaration (widen it, dedup upstream, or drop the claim) rather than to the data.

  • maintain semantic [<objects>] detects definition drift: metric, measure, dimension, or entity definitions that changed against the baseline; semantic references that no longer resolve to a model or column; and categorical dimensions whose set of values widened or narrowed underneath their metrics.

  • maintain reconcile [<class>] proposes the dbt edits that bring the project back in sync, as reviewable diffs. Optionally scope it to one class (schema, volume, grain, or semantic).

The usual flow: check to triage, a focused detector to understand one axis in depth, then reconcile to get the proposed fix.

Per-axis cost: what is free and what scans

Detection is read-only, but read-only is not the same as free on a metered connector (BigQuery, Snowflake, Databricks, Postgres, Redshift, ClickHouse). The axes split:

  • Schema, volume, and the reference/definition half of semantic are free everywhere: they read metadata and the snapshot, and run immediately.
  • Grain and the dimension-cardinality half of semantic scan the warehouse, so on a metered connector they run the two-step handshake. Asked for directly, maintain grain returns needs_confirmation with an estimate in cost.estimate (and a per-table breakdown). Surface it to the user in human units, get an explicit budget, and re-issue the same command with --confirm --budget <magnitude> in the paradigm's unit (bytes on BigQuery, warehouse-seconds on Snowflake and Databricks, compute-seconds on Redshift, database-seconds on Postgres and ClickHouse). Never invent a budget the user did not agree to, and never retry with a raised budget on an over-ceiling refusal without asking. An over-ceiling refusal carries a calibration line from .dex/spend.jsonl (what this connector's recent commands billed as a fraction of estimate, or a sentence saying there is too little history to say): relay it, and note that the ceiling binds on the estimate, so a budget set at that fraction of the estimate is refused again.
  • check and semantic answer first and offer second. Their free axes complete on every call, so the envelope is ok and the findings in it are final. The price of the scanning axes sits in data.offer, with axes naming what it would add; data.axes_run names what already ran. Confirming is a choice, not a required next step: quote the estimate, say which axes are still dark, and let the user decide. A triage pass that stops at the free axes is a complete piece of work, not an abandoned one.
  • Read warnings on these responses, always. They carry the reasons the baseline may no longer describe the warehouse (a cache newer than the snapshot, a baseline pinned from a stale cache), which bound every finding above them. A stale baseline is often the most important line in the response and it is never in findings.

On DuckDB everything is free and local, so nothing prompts.

A needs_confirmation envelope carrying suggested_session_ceiling is the project's one-time ask for a cumulative daily cap, separate from the per-command --budget. Surface it, get the user's answer, and add --session-ceiling <value> or --no-session-ceiling to the same re-issue; it is written to .dex/config.yml once and never asked again. Never answer it for them.

Reconcile proposals are mechanical or advisory

Reconcile tags every proposal by kind, because the fix differs sharply by axis:

  • mechanical: schema drift on a dex-scaffolded staging model re-scaffolds the model from the drifted source. High-confidence, but still a reviewable diff: read it for hand-written logic the scaffold cannot know about.
  • advisory: grain, volume, and semantic drift are decisions, not auto-fixes (dex cannot dedup your warehouse or decide whether a new 'refunded' status belongs in a metric). The proposal is the decision surfaced, at most backed by a test edit that makes the break visible in builds. It declines that test where the test would be wrong: if your model declares a composite grain covering the column, no column-level unique is proposed on it, and the warning names the combination so you can tell "re-baseline, this is still the grain" from "something relied on that column alone".

When reconcile produces edits it stores them as a plan and prints a plan_id. Apply them with transform apply <plan-id> (the one apply door): a human edit made since detection surfaces as a conflict, never a silent overwrite.

Guardrails (enforced in the engine, not here)

  • Read-only against data. Schema, volume, and semantic references are computed from metadata and the snapshot; grain and dimension-cardinality use aggregates only. Raw rows and dimension values never cross the envelope.
  • Propose, don't impose. Reconciliation is always a reviewable diff, applied through transform apply. Human dbt edits are authoritative; on conflict the engine surfaces the divergence and asks rather than overwriting.
  • The dbt project is the source of truth; the .dex/ snapshot is a non-canonical fingerprint used only to detect change.

Signals

GitHub stars
101
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
maintain
Source
github.com/altman-conquer/agentrejoin