cli-eval

SkillDev tools

An AI gateway with enhanced compliance for the domestic market

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the cli-eval skill

What this skill tells your AI

The instructions your AI receives, as published by jaccen/airoute in skills/cli-eval/SKILL.md and read by ahel’s review.

Overview

Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.

Quick install

npm install -g AIRoute   # or: npx AIRoute
AIRoute --version

Subcommands

eval

Example:

AIRoute eval

eval suites

Example:

AIRoute eval suites

eval list

Example:

AIRoute eval list

eval get <suiteId>

Example:

AIRoute eval get <suiteId>

eval create

Flags:

  • --file <path>

Example:

AIRoute eval create

eval run <suiteId>

Flags:

  • -m, --model <id>
  • --combo <name>
  • --concurrency <n>
  • --tag <tag>
  • --watch

Example:

AIRoute eval run <suiteId>

eval list

Flags:

  • --suite <id>
  • --status <s>
  • --since <ts>
  • --limit <n>

Example:

AIRoute eval list

eval get <runId>

Example:

AIRoute eval get <runId>

eval results <runId>

Flags:

  • --failed

Example:

AIRoute eval results <runId>

eval cancel <runId>

Flags:

  • --yes

Example:

AIRoute eval cancel <runId>

eval scorecard <runId>

Example:

AIRoute eval scorecard <runId>

simulate [prompt]

Flags:

  • --file <path>
  • -m, --model <id>
  • --combo <name>
  • --reasoning-effort <level>
  • --thinking-budget <n>
  • --explain

Example:

AIRoute simulate [prompt]

AIRoute — CLI Evals

Requires the AIRoute CLI. See CLI entry-point skill for install + global flags.

What are evals?

Evals are automated test suites that score LLM outputs against expected answers or rubrics. AIRoute stores suites and run results in its local database.

Eval suites

AIRoute eval suites list                       # List all eval suites
AIRoute eval suites list --json                # JSON output

AIRoute eval suites get <suiteId>              # Full suite definition

Create a suite

AIRoute eval suites create \
  --name "code-quality" \
  --rubric "exact-match" \
  --samples-file ./samples.jsonl                 # JSONL: {input, expected_output}

Rubric options: exact-match, contains, llm-judge, regex.

--samples-file format (one JSON object per line):

{"input": "What is 2+2?", "expected_output": "4"}
{"input": "Translate 'hello' to Spanish", "expected_output": "hola"}

Run an eval

AIRoute eval suites run <suiteId> \
  --model claude-sonnet-4-6                      # Run suite against a specific model

AIRoute eval suites run <suiteId> \
  --model gpt-4o \
  --watch                                        # Live TUI progress (EvalWatch)

The run is asynchronous. Use --watch for a live terminal dashboard or poll manually:

RUN_ID=$(AIRoute eval suites run <suiteId> --model claude-sonnet-4-6 --output json | jq -r '.id')
AIRoute eval get $RUN_ID

Manage runs

AIRoute eval list                              # List all eval runs
AIRoute eval list --json

AIRoute eval get <runId>                       # Run details (status, model, score)
AIRoute eval results <runId>                   # Per-sample results
AIRoute eval scorecard <runId>                 # Full scorecard with pass/fail per sample
AIRoute eval cancel <runId>                    # Cancel a running eval

Scorecard output

AIRoute eval scorecard <runId> --output json

Response fields per sample:

{
  "id": "sample-1",
  "score": 0.95,
  "passed": true,
  "input": "What is 2+2?",
  "output": "4",
  "expected": "4"
}

Comparing models

Run the same suite against multiple models and compare:

for MODEL in claude-sonnet-4-6 gpt-4o gemini-2.0-flash; do
  AIRoute eval suites run $SUITE_ID --model $MODEL --output json | jq '{model: .model, score: .score}'
done

CI integration

# Run and fail CI if score drops below threshold
SCORE=$(AIRoute eval suites run $SUITE_ID --model claude-sonnet-4-6 --output json | jq -r '.score')
python3 -c "import sys; score=float('$SCORE'); sys.exit(0 if score >= 0.90 else 1)"

Errors

  • suites create fails with invalid rubric → use one of: exact-match, contains, llm-judge, regex
  • suites run returns model not found → verify model ID with AIRoute models --search <name>
  • eval get shows status: failed → check AIRoute logs --search eval for error details
  • scorecard returns empty results → the run may still be running; poll AIRoute eval get <runId> until status is completed

Signals

GitHub stars
32
Forks
8
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
cli-eval-jaccen
Source
github.com/jaccen/airoute