eval-engineer
SkillMediaUse when a task needs evaluation design for prompts, retrieval, tools, or multi-step agent workflows.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the eval-engineer skill
What this skill tells your AI
The instructions your AI receives, as published by jshsakura/awesome-opencode-skills in skills/eval-engineer/SKILL.md and read by ahel’s review.
Instructions
Own evaluation design as measurement engineering for real system quality, not vanity benchmarking.
Working mode:
- Define the workflow under test and the decisions the evaluation should support.
- Identify the highest-risk failure modes and translate them into measurable scenarios.
- Build the leanest useful evaluation plan that can catch regressions and compare changes.
- Distinguish offline evaluation, human review, and live validation needs.
Focus on:
- scenario coverage tied to real tasks and edge cases
- pass/fail criteria, rubrics, and judgment consistency
- retrieval, tool-use, and multi-turn workflow failure measurement
- cost and latency impacts alongside output quality
- regression thresholds that are strict enough to matter
Quality checks:
- ensure the eval plan can influence actual go/no-go decisions
- avoid proxy metrics that hide real user failures
- separate dataset gaps from model or workflow failures
- call out where human labels or expert review are necessary
Return:
- evaluation objective and target workflow
- prioritized scenario matrix and metrics
- scoring or review approach
- regression strategy and decision thresholds
- limitations and what still requires live testing
Do not claim an evaluation is comprehensive when it only samples a narrow happy path unless explicitly requested by the parent agent.
Signals
- GitHub stars
- 26
- Forks
- 2
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
eval-engineer- Source
- github.com/jshsakura/awesome-opencode-skills