LLM Hallucination Testing
SkillAI & modelsGuides your agent to check its answers against evidence, flag unsupported claims, and state uncertainty instead of guessing.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the LLM Hallucination Testing skill
About this capability
Use this skill when you need evidence-bounded claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review; triggers include LLM 幻觉 and LLM hallucination.
What this skill tells your AI
The instructions your AI receives, as published by naodeng/awesome-qa-skills in skills/en/testing-types/llm-hallucination-testing/SKILL.md and read by ahel’s review.
When to Use
- Use this skill when you need evidence-bounded analysis, design, or validation preparation for claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review.
- Use it to review an Agent, RAG, or LLM plan, result, or evidence set and produce actionable improvements.
- Use it when context is incomplete but a bounded first pass with assumptions, gaps, and human-decision boundaries is still useful.
Output Format Options
- Default to Markdown organized by domain risk, evidence state, priority, and boundary.
- When the user requests tables, CSV, JSON, or ticket fields, preserve the same finding fields, evidence, and decision boundaries.
- Before machine consumption, confirm the schema, enums, required fields, and evidence sources.
How to Use
- Read and follow
prompts/llm-hallucination-testing.md, including its input audit, domain coverage, and output order. - Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to claim-to-source relation, unsupported assertion, abstention, uncertainty, evidence review.
- Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.
- Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.
- When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.
Reference Files
- Always read
prompts/llm-hallucination-testing.md; it is the complete execution specification for this skill. - For evaluation, read
evals/eval.yamland the matching cases underevals/cases/. - Load
references/,examples/,scripts/, oroutput-formats.mdonly when those directories exist and the task needs them.
Core Constraints
- Keep the analysis focused on claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review; do not replace business owners or Human risk acceptance, exception approval, or safety decisions.
- Never invent system behavior, fields, model outputs, data, thresholds, root causes, execution records, or pass claims.
- Static design, plans, file presence, or a dry run retain their evidence state and cannot become proof of real execution.
- When evidence is insufficient, use pending confirmation, blocked, unassessed, or NOT_SCORED and give the smallest validation method.
- For user data, production, or safety work, use least privilege, masked data, mocks, dry runs, or isolation.
Delivery Checklist
- Covered claim-to-source relation, unsupported assertion, abstention, uncertainty, evidence review, with source, evidence state, and validation method for each.
- Separated facts, inferences, candidate recommendations, gaps, and Human decisions.
- Gave high-risk items P0/P1/P2/P3 or an equivalent priority, owner role, and close condition.
- Did not turn plans, static checks, or dry runs into test execution, all-passed, or safety-approved claims.
- Stated residual risk, stop/escalation conditions, and next actions.
Common Pitfalls
- Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.
- Treating adjacent tests or model tools as a complete LLM hallucination judgment.
- Using unexplained numbers for false precision or writing correlation as causation.
- Refusing incomplete input, or pretending that incomplete evidence is conclusive.
Best Practices
- Start with paths most likely to cause user harm, business loss, or decision blockage.
- Use the smallest verifiable experiment to reduce uncertainty and record conditions, versions, sources, and evidence.
- Make the Skill independently installable, executable, and reviewable by another engineer.
Signals
- GitHub stars
- 217
- Forks
- 31
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
llm-hallucination-testing- Source
- github.com/naodeng/awesome-qa-skills