RLHF Self-Reflection Scorer

SkillDocs & knowledge

Offline experimental post-call feedback scorer using supplied ratings and text heuristics. Use to demonstrate advisory review suggestions; no LLM inference, RLHF training, or memory integration is included.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the RLHF Self-Reflection Scorer skill

What this skill tells your AI

The instructions your AI receives, as published by calle-ai/awesome-phone-call-agents in skills/call-rlhf-self-reflection-scorer/SKILL.md and read by ahel’s review.

This skill demonstrates post-call QA with a small lexical scorer. It accepts a supplied transcript and optional rating, then returns predefined review suggestions. It does not collect user feedback, call an LLM, train a model, store RAG memories, or apply prompt changes. These are possible future host integrations, not delivered behavior.

Scientific Foundation

Paper / ConceptRelevance
LLM-as-a-JudgeUsing a strong LLM to evaluate the outputs of an agentic LLM correlates highly with human CSAT (Customer Satisfaction).
Self-Reflection (Reflexion)Agents that critique their own past transcripts and generate "verbal reinforcement" prompts perform significantly better on subsequent tasks.
RLHF (Reinforcement Learning from Human Feedback)Incorporating explicit user scores (if provided post-call) alongside automated critiques bridges the gap between simulated and real-world quality.

How it works

  1. The skill receives the transcript and any explicit CSAT score given by the user (if applicable).
  2. If the user gave a high score (>= 4), the interaction is marked as successful.
  3. If the score is low or missing, the skill checks a small set of text patterns and returns a mocked critique; it cannot establish a root cause.
  4. It outputs an EvaluationResult containing the score, the identified critique, and a specific system prompt recommendation to fix the behavior.

Decision Matrix

Explicit ScoreTranscript SentimentOutcomeAction
>= 4AnyEXPLICIT_USER (High)Maintain current strategy
< 4AnySELF_CRITIQUEReturn a heuristic critique; not an explanation of the user's rating
NoneSmoothSELF_CRITIQUE (High)Baseline evaluation
NoneFriction detectedSELF_CRITIQUE (Low)Flag friction point and generate patch

Expected Outcomes & Metrics

These are unvalidated targets for a possible future evaluator, not measurements of this mock scorer.

MetricTargetNotes
Critique Relevance> 90%The LLM-generated critique should match human QA audits.
Recommendation Actionability> 85%Recommendations must be directly usable as system prompt instructions.

Limitations & Known Constraints

  • Self-Correction Loop: This skill only generates the critique. A separate meta-agent is required to actually update the core agent's prompt based on these recommendations.
  • Cost: The local helper has no LLM dependency. A future LLM integration would have separate costs and evaluation needs.

Signals

GitHub stars
104
Forks
527
Last commit
Sep 2026
Advanced
Item type
skill
Key
call-rlhf-self-reflection-scorer
Source
github.com/calle-ai/awesome-phone-call-agents