agent-evaluation
SkillMonitoring & opsTesting and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarksUse when "agent testing, agent evaluation, benchmark agents, agent reliability, test agent, test
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the agent-evaluation skill
Signals
- GitHub stars
- 144
- Forks
- 22
- Last commit
- Jan 2026
Advanced
- Catalog kind
- skill
- Gateway key
agent-evaluation-omer-metin- Source
- github.com/omer-metin/skills-for-antigravity