Benchmark — Performance Baseline & Regression Detection

SkillDev tools

Benchmark lets your AI measure how your pages, APIs, and builds perform and compare the numbers before and after code changes. That way slowdowns and regressions get caught early instead of after they ship. It can also compare stack alternatives so you can choose the option that performs best.

Available today. Use it from your connected AI after setup.

After adding it, ask your AI to record a baseline for the pages, APIs, or builds you care about. From then on, it can compare new measurements against that baseline whenever your code changes.

Then ask your AI: use the Benchmark — Performance Baseline & Regression Detection skill

What your AI can do with it

  • Measure page, API, and build performance
  • Set performance baselines for your project
  • Detect regressions by comparing results before and after code changes
  • Spot regressions around pull requests
  • Compare stack alternatives on performance

What this skill tells your AI

The instructions your AI receives, as published by affaan-m/ecc in skills/benchmark/SKILL.md and read by ahel’s review.

When to Use

  • Before and after a PR to measure performance impact
  • Setting up performance baselines for a project
  • When users report "it feels slow"
  • Before a launch — ensure you meet performance targets
  • Comparing your stack against alternatives

How It Works

Mode 1: Page Performance

Measures real browser metrics via browser MCP:

1. Navigate to each target URL
2. Measure Core Web Vitals:
   - LCP (Largest Contentful Paint) — target < 2.5s
   - CLS (Cumulative Layout Shift) — target < 0.1
   - INP (Interaction to Next Paint) — target < 200ms
   - FCP (First Contentful Paint) — target < 1.8s
   - TTFB (Time to First Byte) — target < 800ms
3. Measure resource sizes:
   - Total page weight (target < 1MB)
   - JS bundle size (target < 200KB gzipped)
   - CSS size
   - Image weight
   - Third-party script weight
4. Count network requests
5. Check for render-blocking resources

Mode 2: API Performance

Benchmarks API endpoints:

1. Hit each endpoint 100 times
2. Measure: p50, p95, p99 latency
3. Track: response size, status codes
4. Test under load: 10 concurrent requests
5. Compare against SLA targets

Mode 3: Build Performance

Measures development feedback loop:

1. Cold build time
2. Hot reload time (HMR)
3. Test suite duration
4. TypeScript check time
5. Lint time
6. Docker build time

Mode 4: Before/After Comparison

Run before and after a change to measure impact:

/benchmark baseline    # saves current metrics
# ... make changes ...
/benchmark compare     # compares against baseline

Output:

| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |

Output

Stores baselines in .ecc/benchmarks/ as JSON. Git-tracked so the team shares baselines.

Integration

  • CI: run /benchmark compare on every PR
  • Pair with /canary-watch for post-deploy monitoring
  • Pair with /browser-qa for full pre-ship checklist

Signals

GitHub stars
256k
Forks
38k
Last commit
Sep 2026
Hacker News mentions
20
Advanced
Catalog kind
skill
Gateway key
benchmark
Source
github.com/affaan-m/ecc