prompt-testing

SkillMonitoring & ops

Use when comparing two prompt variants, defining quality/efficiency/robustness metrics, or deciding whether to adopt a challenger prompt over a baseline.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the prompt-testing skill

What this skill tells your AI

The instructions your AI receives, as published by fusengine/agents in plugins/prompt-engineer/skills/prompt-testing/SKILL.md and read by ahel’s review.

The adoption decision is rule-based: adopt B if its accuracy is at least equal with acceptable token cost, consider B as a trade-off if accuracy improves >10% despite <20% token regression, otherwise keep A or iterate. Requires a minimum of 20 test cases with 15-20% edge cases for statistical significance.

Prompt Testing

Skill for testing, comparing, and measuring prompt performance.

References

  • metrics.md - Load when: defining or scoring Quality/Efficiency/Robustness/UX metrics with thresholds and calculation formulas
  • methodology.md - Load when: running a full A/B test (hypothesis, dataset sizing, statistical significance, common pitfalls)
  • templates.md - Load when: writing a test dataset JSON or an A/B test report

Testing Workflow

1. DEFINE
   └── Test objective
   └── Metrics to measure
   └── Success criteria

2. PREPARE
   └── Variants A and B
   └── Test dataset
   └── Baseline (if existing)

3. EXECUTE
   └── Run on dataset
   └── Collect results
   └── Document observations

4. ANALYZE
   └── Calculate metrics
   └── Compare variants
   └── Identify patterns

5. DECIDE
   └── Recommendation
   └── Statistical confidence
   └── Next iterations

Performance Metrics

Quality

MetricDescriptionCalculation
AccuracyCorrect responsesCorrect / Total
ComplianceFormat adherenceCompliant / Total
ConsistencyResponse stability1 - Variance
RelevanceMeeting the needAverage score (1-5)

Efficiency

MetricDescriptionCalculation
Tokens InputPrompt sizeToken count
Tokens OutputResponse sizeToken count
LatencyResponse timems
CostPrice per requestTokens × Price

Robustness

MetricDescriptionCalculation
Edge CasesEdge case handlingPassed / Total
Jailbreak ResistBypass resistanceBlocked / Attempts
Error RecoveryError recoveryRecovered / Errors

For full definitions, thresholds, and the UX metrics category, see metrics.md. For the test dataset and report formats, see templates.md.

Commands

# Create a test
/prompt test create --name "Test v1" --dataset tests.json

# Run an A/B test
/prompt test run --a prompt_a.md --b prompt_b.md --dataset tests.json

# View results
/prompt test results --id test_001

# Compare two tests
/prompt test compare --tests test_001,test_002

Decision Criteria

When to adopt variant B?

IF:
  - Accuracy B >= Accuracy A
  AND (Tokens B <= Tokens A * 1.1 OR accuracy improvement > 5%)
  AND no regression on edge cases
THEN:
  → Adopt B

ELSE IF:
  - Accuracy improvement > 10%
  AND token regression < 20%
THEN:
  → Consider B (acceptable trade-off)

ELSE:
  → Keep A or iterate

Best Practices

  1. Minimum 20 test cases for significance
  2. Include edge cases (15-20% of dataset)
  3. Test multiple runs for consistency
  4. Document hypotheses before testing
  5. Version the prompts being tested

Signals

GitHub stars
25
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
prompt-testing
Source
github.com/fusengine/agents