memoriant-eval-sandbox-skill

PackDev tools

Lets your agent test other agents against holdout scenarios and score their performance before deployment.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

AI agent evaluation sandbox for Claude Code. Holdout scenario testing with Doer/Judge/Adversary/Observer roles, probabilistic satisfaction scoring, and append-only JSONL audit trails with integrity hashes. Test your agents before deploying them. Includes full Python evaluation framework.

Signals

GitHub stars
4k
Forks
312
Last commit
Aug 2026
Advanced
Item type
plugin
Key
anthropics-claude-plugins-community-memoriant-eval-s-1ai5io0
Source
github.com/anthropics/claude-plugins-community