promptfoo-evaluation

PackAI & models

Lets your agent set up and run tests that compare AI model outputs against expected results.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

Configures and runs LLM evaluation using Promptfoo framework. Use when setting up prompt testing, creating evaluation configs (promptfooconfig.yaml), writing Python custom assertions, implementing llm-rubric for LLM-as-judge, or managing few-shot examples in prompts. Triggers on keywords like prompt

Signals

GitHub stars
1k
Forks
220
Last commit
Sep 2026
Advanced
Item type
plugin
Key
daymade-claude-code-skills-promptfoo-evaluation
Source
github.com/daymade/claude-code-skills