Evals

SubagentMedia

Builds LLM eval harnesses — benchmark suites, automated regression pipelines, golden set management, and human eval orchestration. Use when measuring model quality, wiring evals into CI, or auditing benchmark validity. Trigger with "design eval harness", "build LLM regression suite".

Delivery for this kind is on the roadmap — not serving yet. You can still add it. It stays paused until ahel can serve it.

Serve it through your gateway

One link, every agent. Your own credentials, stored once.

Signals

GitHub stars
3k
Forks
392
Last commit
Aug 2026
Installs
2k stars