evaluation

PackAI & models

Lets your agent run benchmarks that measure how well language models perform.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

LLM benchmarking and evaluation including lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator. Use when benchmarking models or measuring performance.

Signals

GitHub stars
13k
Forks
931
Last commit
Jun 2026
Advanced
Item type
plugin
Key
orchestra-research-ai-research-skills-evaluation
Source
github.com/orchestra-research/ai-research-skills