Evals
SubagentMediaBuilds LLM eval harnesses — benchmark suites, automated regression pipelines, golden set management, and human eval orchestration. Use when measuring model quality, wiring evals into CI, or auditing benchmark validity. Trigger with "design eval harness", "build LLM regression suite".
Delivery for this kind is on the roadmap — not serving yet. You can still add it. It stays paused until ahel can serve it.
Serve it through your gateway
One link, every agent. Your own credentials, stored once.
Signals
- GitHub stars
- 3k
- Forks
- 392
- Last commit
- Aug 2026
- Installs
- 2k stars