A/B test agent skills and MCP changes with Caliper

SkillProductivity

Use Caliper to run real agent tasks with and without a skill, MCP server, or rule change so reliability and token cost are measurable.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the A/B test agent skills and MCP changes with Caliper skill

What this skill tells your AI

The instructions your AI receives, as published by agentskillexchange/skills in skills/ab-test-agent-skills-mcp-changes-with-caliper/SKILL.md and read by ahel’s review.

Use Caliper to run real agent tasks with and without a skill, MCP server, or rule change so reliability and token cost are measurable.

Prerequisites

Python 3.10+, Caliper CLI, target agent runtime such as Claude Code, Codex, Pi, or Hermes, skills or MCP servers under test

Installation

Install or set up from the source-backed instructions:

Install the CLI with pipx install caliper-eval for direct runs, or add the agent skill with npx skills@latest add edonadei/caliper; then create an .eval.yaml, run caliper run ... --k 3, run an ablated control with --ablate, and compare results with caliper compare.

Documentation

Source

Signals

GitHub stars
42
Forks
54
Last commit
Sep 2026
Advanced
Catalog kind
skill
Key
a-b-test-agent-skills-and-mcp-changes-with-caliper
Source
github.com/agentskillexchange/skills