evaluation
PackAI & modelsLets your agent run benchmarks that measure how well language models perform.
Unavailable. Delivery for this kind is on the roadmap — not serving yet.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this app
LLM benchmarking and evaluation including lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator. Use when benchmarking models or measuring performance.
Signals
- GitHub stars
- 13k
- Forks
- 931
- Last commit
- Jun 2026
Advanced
- Item type
- plugin
- Key
orchestra-research-ai-research-skills-evaluation- Source
- github.com/orchestra-research/ai-research-skills