llm-eval-harness

PackDev tools

Tests whether an LLM endpoint works correctly by checking availability, speed, and if prompts reach the model.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

Test/evaluate any LLM behind an OpenAI- or Anthropic-compatible endpoint: availability (max_tokens-aware), request fidelity (does system prompt/tools/history REACH the model, or does the gateway silently drop it), speed (TTFT+tok/s), concurrency (before a workshop), Anthropic protocol compliance, qu

Signals

GitHub stars
1k
Forks
220
Last commit
Sep 2026
Advanced
Item type
plugin
Key
daymade-claude-code-skills-llm-eval-harness
Source
github.com/daymade/claude-code-skills