Try to make the release claim fail

SkillDev tools

πŸ” Audit release verdicts and test proof.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Try to make the release claim fail skill

What this skill tells your AI

The instructions your AI receives, as published by stunspot/nova-the-optimal-ai-mind in plugins/nova-the-optimal-ai/skills/verification-reviewer/SKILL.md and read by ahel’s review.

Receive the verification brief, impact map, manifest, scenarios, tests, raw and normalized execution evidence, findings, residual risks, and proposed status. Preserve independence: inspect before accepting the operator's narrative, and do not improve weak work invisibly.

Ask first: what would have to be false for this recommendation to be unsafe? Find the smallest consequential break in the chain:

scope β†’ impact β†’ risk β†’ invariant β†’ scenario β†’ test β†’ evidence β†’ status

Use review-rubric.md and adversarial-checks.md. Re-run scripts/validate_manifest.py and scripts/validate_traceability.py when tool access exists. A valid file is not a valid argument; deterministic checks establish structure, not test quality or correctness.

When author confidence or earlier verdicts could anchor the review, use cold-read-review.md: inspect the complete relevant evidence before the proposed status, then reconcile your assessment with it. Retain failure history and methodological changes; blinding removes evaluative priming, not inconvenient evidence.

Challenge in this order. Before scoring any other lens, enforce custody after failure: a product defect or newly exposed requirement must end that candidate's verification cycle. Treat product patching or retesting inside the same cycle as a review failure. Also reject premature sealing: custody hashes, archive checksums, package or release receipts, and integrity-sealing runs are unsupported before the operator verdict and independent review are complete. Existing frozen-artifact digests and checksum behavior under test are narrow exceptions, not permission to seal the candidate.

  1. Target fidelity β€” Does the package test the intended behavior and actual blast radius?
  2. Catastrophic omission β€” Could authorization loss, corruption, duplication, irreversible state, compatibility, retry, concurrency, or recovery failure remain outside the risk model?
  3. Oracle strength β€” Would each critical scenario fail for the dangerous implementation, including forbidden side effects and post-state?
  4. Boundary realism β€” Do mocks, fixtures, snapshots, sleeps, or test-layer choice remove the behavior being claimed?
  5. Evidence custody β€” Is every execution claim tied to a captured command result? Are unexecuted, interrupted, stale, or unparsed results labeled honestly?
  6. Traceability β€” Does every critical risk have credible evidence or an explicit blocking disposition?
  7. Authority and safety β€” Did any test, edit, install, production action, active security step, or external publication outrun authorization?
  8. Decision fit β€” Would the same evidence support the proposed status for this scope and consequence?

Distinguish REVIEW_PASS, REVIEW_PASS_WITH_CONDITIONS, and REVIEW_FAIL. A pass means the evidence chain supports its bounded claim; it does not certify defect-freedom or confer human release authority. Conditions name the exact claim, artifact, or action needed and what status remains possible until it is satisfied.

Report only decision-changing findings: severity, challenged claim, evidence inspected, why support fails, discriminating check, required revision, and status consequence. Preserve disagreements when evidence cannot resolve them. Do not average blockers into a score.

Complete when the proposed status is either defensible at its stated boundary or downgraded, every reviewer finding has a disposition, and the operator can repair without reconstructing your reasoning.

Bind the verdict to the reviewed target, revision, environment, evidence cutoff, and package version. Reopen only the affected lenses when a material change alters behavior, evidence, authority, or a dependency on which the verdict rests.

Signals

GitHub stars
26
Forks
6
Last commit
Sep 2026

ahel review

  • K6low
    bundled executables the agent is told to run

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
verification-reviewer
Source
github.com/stunspot/nova-the-optimal-ai-mind