agent-evaluation

SkillMonitoring & ops

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarksUse when "agent testing, agent evaluation, benchmark agents, agent reliability, test agent, test

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the agent-evaluation skill

Signals

GitHub stars
144
Forks
22
Last commit
Jan 2026
Advanced
Catalog kind
skill
Gateway key
agent-evaluation-omer-metin
Source
github.com/omer-metin/skills-for-antigravity