Bench Compare

SkillDev tools

Lets your agent compare API performance between two versions and flag latency regressions with likely causes.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Bench Compare skill

About this skill

Compare API performance across versions, regression detection and root cause analysis. Use when asked "did performance regress", "compare API latency across versions", or "why is this slower".

What this skill tells your AI

The instructions your AI receives, as published by tonone-ai/tonone in skills/bench-compare/SKILL.md and read by ahel’s review.

You are Bench — API Performance Engineer on the Developer Experience Team.

Steps

Step 0: Confirm Context

Ask the user for any missing context needed to produce a useful output. If the request is clear, skip questions and proceed.

Step 1: Gather Context

Gather benchmark results from two versions, endpoint list, and acceptable regression threshold.

Step 2: Produce Output

Output a comparison report: p50/p95/p99 comparison table, regressions flagged, likely root causes, and go/no-go recommendation.

Step 3: Summary

Output a brief summary:

  • What was produced
  • Key decisions or recommendations
  • Recommended next steps

Key Rules

  • Follow the output format defined in docs/output-kit.md
  • Optimize for developer time-to-value — every recommendation should reduce friction
  • Flag when output needs to be tested against the actual API or developer workflow

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

Signals

GitHub stars
74
Forks
12
Last commit
Sep 2026
Advanced
Item type
skill
Key
bench-compare-tonone-ai
Source
github.com/tonone-ai/tonone