memoriant-eval-sandbox-skill
PackDev toolsLets your agent test other agents against holdout scenarios and score their performance before deployment.
Unavailable. Delivery for this kind is on the roadmap — not serving yet.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this app
AI agent evaluation sandbox for Claude Code. Holdout scenario testing with Doer/Judge/Adversary/Observer roles, probabilistic satisfaction scoring, and append-only JSONL audit trails with integrity hashes. Test your agents before deploying them. Includes full Python evaluation framework.
Signals
- GitHub stars
- 4k
- Forks
- 312
- Last commit
- Aug 2026
Advanced
- Item type
- plugin
- Key
anthropics-claude-plugins-community-memoriant-eval-s-1ai5io0- Source
- github.com/anthropics/claude-plugins-community