cli-eval
SkillDev toolsAn AI gateway with enhanced compliance for the domestic market
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the cli-eval skill
What this skill tells your AI
The instructions your AI receives, as published by jaccen/airoute in skills/cli-eval/SKILL.md and read by ahel’s review.
Overview
Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.
Quick install
npm install -g AIRoute # or: npx AIRoute
AIRoute --version
Subcommands
eval
Example:
AIRoute eval
eval suites
Example:
AIRoute eval suites
eval list
Example:
AIRoute eval list
eval get <suiteId>
Example:
AIRoute eval get <suiteId>
eval create
Flags:
--file <path>
Example:
AIRoute eval create
eval run <suiteId>
Flags:
-m, --model <id>--combo <name>--concurrency <n>--tag <tag>--watch
Example:
AIRoute eval run <suiteId>
eval list
Flags:
--suite <id>--status <s>--since <ts>--limit <n>
Example:
AIRoute eval list
eval get <runId>
Example:
AIRoute eval get <runId>
eval results <runId>
Flags:
--failed
Example:
AIRoute eval results <runId>
eval cancel <runId>
Flags:
--yes
Example:
AIRoute eval cancel <runId>
eval scorecard <runId>
Example:
AIRoute eval scorecard <runId>
simulate [prompt]
Flags:
--file <path>-m, --model <id>--combo <name>--reasoning-effort <level>--thinking-budget <n>--explain
Example:
AIRoute simulate [prompt]
AIRoute — CLI Evals
Requires the AIRoute CLI. See CLI entry-point skill for install + global flags.
What are evals?
Evals are automated test suites that score LLM outputs against expected answers or rubrics. AIRoute stores suites and run results in its local database.
Eval suites
AIRoute eval suites list # List all eval suites
AIRoute eval suites list --json # JSON output
AIRoute eval suites get <suiteId> # Full suite definition
Create a suite
AIRoute eval suites create \
--name "code-quality" \
--rubric "exact-match" \
--samples-file ./samples.jsonl # JSONL: {input, expected_output}
Rubric options: exact-match, contains, llm-judge, regex.
--samples-file format (one JSON object per line):
{"input": "What is 2+2?", "expected_output": "4"}
{"input": "Translate 'hello' to Spanish", "expected_output": "hola"}
Run an eval
AIRoute eval suites run <suiteId> \
--model claude-sonnet-4-6 # Run suite against a specific model
AIRoute eval suites run <suiteId> \
--model gpt-4o \
--watch # Live TUI progress (EvalWatch)
The run is asynchronous. Use --watch for a live terminal dashboard or poll manually:
RUN_ID=$(AIRoute eval suites run <suiteId> --model claude-sonnet-4-6 --output json | jq -r '.id')
AIRoute eval get $RUN_ID
Manage runs
AIRoute eval list # List all eval runs
AIRoute eval list --json
AIRoute eval get <runId> # Run details (status, model, score)
AIRoute eval results <runId> # Per-sample results
AIRoute eval scorecard <runId> # Full scorecard with pass/fail per sample
AIRoute eval cancel <runId> # Cancel a running eval
Scorecard output
AIRoute eval scorecard <runId> --output json
Response fields per sample:
{
"id": "sample-1",
"score": 0.95,
"passed": true,
"input": "What is 2+2?",
"output": "4",
"expected": "4"
}
Comparing models
Run the same suite against multiple models and compare:
for MODEL in claude-sonnet-4-6 gpt-4o gemini-2.0-flash; do
AIRoute eval suites run $SUITE_ID --model $MODEL --output json | jq '{model: .model, score: .score}'
done
CI integration
# Run and fail CI if score drops below threshold
SCORE=$(AIRoute eval suites run $SUITE_ID --model claude-sonnet-4-6 --output json | jq -r '.score')
python3 -c "import sys; score=float('$SCORE'); sys.exit(0 if score >= 0.90 else 1)"
Errors
suites createfails withinvalid rubric→ use one of:exact-match,contains,llm-judge,regexsuites runreturnsmodel not found→ verify model ID withAIRoute models --search <name>eval getshowsstatus: failed→ checkAIRoute logs --search evalfor error detailsscorecardreturns empty results → the run may still berunning; pollAIRoute eval get <runId>untilstatusiscompleted
Signals
- GitHub stars
- 32
- Forks
- 8
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
cli-eval-jaccen- Source
- github.com/jaccen/airoute