Anvia Evals Skill
SkillMonitoring & opsEvaluate Anvia agents and retrieval, deterministic metrics, semantic similarity, LLM judges, RAG quality, and CLI eval runs.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Anvia Evals Skill skill
What this skill tells your AI
The instructions your AI receives, as published by anvia-hq/anvia in skills/anvia-evals/SKILL.md and read by ahel’s review.
Use this skill when the user wants to measure output quality: scoring a target function or agent over cases, picking metrics, adding an LLM judge, checking RAG grounding, or wiring evals into CI.
Process
- Start deterministic (
references/metrics.md) — exact match, contains, semantic similarity. - Add judges only for what strings cannot check (
references/judges.md). - Run it right (
references/running.md) —runEvalSuitevsrunEvalCli, negative controls, expectations. - Run
scripts/check-evals.shfrom the app root before claiming done.
Minimal slice
import { contains, exactMatch, runEvalCli } from "@anvia/core/evals";
await runEvalCli({
name: "support-basic-metrics",
cases: [
{
id: "refund-window",
input: "When can I request a refund?",
expected: "Refunds are available for 30 days.",
},
{
id: "wrong-refund-window",
input: "Negative control: when can I request a refund?",
expected: "Refunds are available for 30 days.",
},
],
target: async (input) => answerSupportQuestion(input),
metrics: [exactMatch(), contains({ expected: ({ case: testCase }) => "30 days" })],
expectations: {
outcomes: { "wrong-refund-window": { exact_match: "fail", contains: "fail" } },
},
exitCode: true,
});
Output
Every behavior claim needs a case. Prefer cheap deterministic metrics; spend LLM-judge budget where wording varies. Point to the relevant reference file instead of pasting its contents into chat.
Signals
- GitHub stars
- 50
- Forks
- 6
- Last commit
- Oct 2026
ahel review
K6low
bundled executables the agent is told to run
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
anvia-evals- Source
- github.com/anvia-hq/anvia
Related picks
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptinternal-comms
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opsarchitecture-decision-records
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & ops