ahel is live on Product Hunt today. Upvote

Veris

MCP serverAI & models

Behavioral verification intelligence for AI coding agents. 17 MCP tools. Local-first. MIT.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Connect ahel once, and every AI you use reads what you have installed.

From the project's README

As published by vighriday/veris in README.md.


We pointed Veris at Veris — it found 55 defects

Every one is published — what broke, why it mattered, the fix, and the test that proves it: docs/internal/BUG_TRACKER.md

The worst three, in our own tool:

It invented baselines. When git was unavailable, Veris built a "before" state from the first 70% of the current graph and reported the comparison as a real behavioral diff. No flag. No warning. A verification tool was fabricating the thing it verified against.

91% of its call edges were guesses. It matched the trailing name of a call against every declaration sharing that name. console.log() drew an edge to the project's own Logger.log. Measured on a real dependency: 2,804 of 3,077 edges pointed at an ambiguous name.

The graded agent could erase its own failures. Execution results were stored with INSERT OR REPLACE. Post fail, then post pass, and the failure was gone.

We could have fixed these quietly. Publishing them is the point: a tool that tells you what is unverified has no standing to hide its own unverified claims.

This is also the demo. That is the analysis Veris performs, run on itself.


What Veris is

A behavioral diff for AI-written code, speaking the Model Context Protocol so your agent can ask while it is still working — not after you find out in review.

It answers two questions a line diff cannot:

  1. What behavior changed? Not which lines — which behaviors, and what reaches them.
  2. Was any of it actually checked? Published research puts roughly 65% of agent-authored PRs at zero coverage of their own changed lines.

Veris never executes anything. No tests, no sandboxes, no runtime. It reads, models, and tells your agent what is at risk and what evidence exists. Running things stays with the tools that are good at running things.


The 30-second version

$ npx veris-core . --base-ref=origin/main

-> Baseline: origin/main @ 1bebd2ce2e08 -> head 3ed9031421-dirty
   Working tree has 4 uncommitted changes; this run is not reproducible from commits alone.
-> Graph: 326 nodes, 602 edges (head), 131 tracked files
-> Call resolution: 403 resolved (97.1%), 6 single-candidate, 6 ambiguous (no edge emitted)
-> Workflows: 15 detected, 3 affected in diff
-> Adversarial probes generated: 4

Read lines 2 and 4 again — they are the whole philosophy.

Six calls were too ambiguous to resolve, so Veris drew no edge rather than guessing. The head is marked -dirty because uncommitted changes were included, so the result is not reproducible from commits alone.

Most tools report only what they found. Veris also reports what it could not determine, because a confident wrong answer is worse than an admitted gap.


Install

As an MCP server — one config block, then restart your client:

{
  "mcpServers": {
    "veris": {
      "command": "npx",
      "args": ["-y", "veris-core", "mcp"]
    }
  }
}

17 tools light up in Claude Code, Cursor, or any MCP-compatible agent.

As a CLI:

npx veris-core .                            # analyze against origin/main
npx veris-core . --base-ref=HEAD~1          # explicit baseline
npx veris-core . --budget=10 --onboarding   # 10-min plan + onboarding map
npx veris-core doctor                       # check git, base ref, deps

Needs a git repository with real history. Veris diffs against the merge-base with your base ref. If it cannot establish one, it fails and says why rather than inventing a baseline. In CI: fetch-depth: 0.

On npm 12, run history needs one extra line. npm 12 no longer runs dependency install scripts by default, so better-sqlite3 never fetches its prebuilt binding. Veris still analyzes, diffs, scores risk and plans verification — only run history and cross-run drift need it. The allowlist is per-project and is not inherited from a dependency, so it has to go in your package.json:

{ "allowScripts": { "better-sqlite3": true } }

Then npm rebuild better-sqlite3. veris-core doctor reports which mode you are in, and never claims persistence is working when it is not.


How it thinks

flowchart TD
    A[git-tracked source] -->|ts-morph + TypeScript checker| B[Behavioral graph]
    B -->|worktree at merge-base| C{Baseline exists?}
    C -->|no| X[Fail loudly<br/>never fabricate]
    C -->|yes| D[Diff: added / removed<br/>rewritten-body / edges]
    D --> E[Risk · Workflows · Fingerprints · Drift]
    E --> F[Probes · Tiered plan · Budget]
    F --> G[Coverage from<br/>trust-weighted evidence]
    G --> H[17 MCP tools · Dashboard · Reports]
    H -->|agent or CI executes| I[report_execution]
    I -->|append-only, hash-chained| G

    style X fill:#ff5d6c,stroke:#c1121f,color:#fff
    style G fill:#8b5cf6,stroke:#6d28d9,color:#fff
    style B fill:#0ea5e9,stroke:#0369a1,color:#fff

The red box is a feature. So is the loop back into coverage.


Three ideas that make it different

1. Every edge declares how certain it is

Most graph tools give you an edge. Veris tells you why it believes the edge:

resolutionMeaning
resolvedThe TypeScript checker identified the declaration. Trustworthy.
heuristicChecker couldn't, but exactly one declaration bears that name.
structuralContainment or an import relationship.
(no edge)Several candidates and nothing distinguishes them. Silence, not a guess.

Anything that must not reason on a guess — a gate, a policy rule — filters for resolved. Missing edges understate coupling. They never invent it.

2. Evidence is append-only, and knows who said it

The agent posting results is usually the agent being judged. So:

{ "nodeId": "src/pay.ts::charge",
  "result": "pass",
  "trustClass": "harness-observed",   // ← default is "agent-asserted"
  "producer": "github-actions:e2e" }
Trust classWhoWeight
veris-derivedVeris computed itfull
harness-observedAn external runner saw itfull
agent-assertedThe agent says so — the defaulthalf

Records are hash-chained. A later pass never overwrites an earlier failure; editing the database directly breaks the chain and verifyEvidenceChain() reports exactly where. An agent cannot raise its own assurance by asserting harder.

3. It catches the rewrite that keeps its name

- function chargeCard(amount) { return gateway.charge(amount); }
+ function chargeCard(amount) { return gateway.charge(amount * 100); }

Same name. Same callees. Same graph shape. Every name-and-topology comparison sees nothing. Veris hashes the normalized body, so this surfaces as a modifiedNode — while renaming a directory, which used to look like 100% drift, now correctly looks like nothing at all.


What your agent asks

veris: analyze_pr_behavior with baseRef=origin/main
veris: list_workflows, then analyze_workflow for the highest-risk one
veris: generate_adversarial_probes, then allocate_budget minutes=15
veris: detect_drift
veris: what_if_revert nodeIds=[...]

Probes are concrete, not nudges:

Payments / idempotency — Submit a charge twice with the same idempotency key inside a 500 ms window. Invariant: exactly one ledger entry; the second call returns the first result.


Webhooks / replay — Replay a 24-hour-old signed payload with its original signature. Invariant: rejected by timestamp window even though the signature is valid.


Everything else it does

Semantic workflows25 domains — Authentication, Payments, Checkout, Webhooks, Queue, Caching… So the unit is "checkout reliability", not GraphModels.ts.
Risk modelCoupling magnitude, inbound-coupling dominance, runtime criticality — three inputs measuring different things. Every weight in data/risk-config.json, plain-English reasons attached.
Drift detectionFingerprints across runs. Catches silent rewrites, surface changes, oscillating refactors, and deletions.
Budget allocationGiven N minutes, the highest-leverage subset to actually run.
Counterfactualwhat_if_revert — what recovers if this comes out?
Onboarding exportWorkflow-first markdown for a new engineer, or a new agent, on an unfamiliar codebase.
DashboardStandalone HTML. Click a workflow, everything filters. Click-to-copy directives.

Honest limits

Stated plainly, so nobody discovers them the hard way.

  • A workflow is a label, not a path. Classification is a weighted keyword vote over directory names, imports and symbol names. It does not traverse the call graph. Rate-limiting code that imports Redis lands in Caching. Making workflows real paths is the top roadmap item.
  • Coverage is not assurance. It measures how much planned verification has evidence behind it. It is not calibrated against real incidents and does not estimate the probability your code is correct.
  • Risk is a heuristic. Good for ranking what to look at first. Not a defect predictor. No ground truth behind it.
  • Probes are a curated library — real failure modes, written by hand, selected by domain. Not generated from your code.
  • TypeScript and JavaScript only. Python and Go are on the roadmap.
  • Some calls can't be resolved. Dynamic dispatch and untyped JS defeat the checker. Those produce no edge, and the count is in the output.

Upgrading from 2.x? 3.0 has real breaking changes — see UPGRADING.md.


Privacy & security

  • Local-first. All analysis runs on your machine. No telemetry, ever. Nothing about your code leaves the machine.
  • Zero-retention modeVERIS_STATE_DISABLED=1.
  • No network sockets in the analyzer. stdio and the filesystem only.

Veris is usually pointed at repositories you did not write, so repository content is untrusted input. Plugins execute code from the analyzed repo, so they are off by default--allow-plugins opts in, and each plugin's path and SHA-256 is printed before it runs. There is no sandbox, and SECURITY.md says so plainly instead of implying otherwise.


Docs

MCP toolsAll 17 tools with recommended flows
ArchitectureDesign invariants and the defect each replaced
Audit trackerAll 55 findings, with evidence
Upgrading2.x → 3.0
SecurityThreat model and reporting
RoadmapWhat is next — and what will never be built
PluginsExtending classification and risk

Contributing

The five things that move the needle most:

  1. Entry-point detection for a framework you know — routes, handlers, queue consumers. This is what turns a workflow from a label into a path.
  2. Labelled repositories for a classification benchmark. The accuracy claim needs ground truth, not more rules.
  3. Probe provenance. The shipped probes are good and uncited; one backed by a public postmortem is worth ten that aren't.
  4. Language adapters — Python, Go.
  5. Calibration data — what Veris flagged that broke, and what it missed. The second is more valuable.

See CONTRIBUTING.md. Open source, sponsor-supported. No paid tier, no gated features, no open-core bait.

Signals

GitHub stars
1
Last commit
Sep 2026
Weekly downloads
62
Advanced
Delivery
veris MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
io-github-vighriday-veris
Source
github.com/vighriday/veris